Skip to content
Gridiron

walk-forward, weekly, expanding window

Historical backtest

Accuracy

5,762 games scored from 2005 · 3-season warm-up · de-vig shin · published Oct 6, 2026

Published forecast record · 2026

Decided games
65
Brier
0.2316
Accuracy
61.5%
Pending
207

Earliest recorded pregame forecast per fixture. This small live sample is separate from the historical backtest below. Published Oct 10, 2026.

Break down by model and forecast horizon

model version

CohortGamesBrierLog loss
nfl-margin-lattice-1650.23160.6542

horizon

CohortGamesBrierLog loss
under 24h0——
1 to 7 days0——
7 days or more650.23160.6542

Retained publications · 2026

Small sample

First vs latest forecasts

65 identical decided games · probability scores, lower is better.

First publication

Paired Brier
0.23158
Log loss
0.65415
Games
65

Latest pregame

Paired Brier
0.23119
Log loss
0.65083
Games
65

Latest − first Brier: -0.00039 · descriptive difference only.

Latest under 24h
63 / 65
Missing latest
0
Ties excluded
0

2 decided games have no valid stored forecast under 24 hours before kickoff. This sample does not establish an accuracy gain or support model promotion.

All cohorts, horizons and source coverage

All-cohort scores below can have different game IDs. Only the paired cards above compare identical decided games; ties are counted and excluded.

Separate first-publication and latest-pregame cohorts, with probability scores by horizon and model
Publication / cohortGamesBrierLog loss
First · All decided650.231580.65415
First · Under 24 hours0——
First · 1–7 days0——
First · 7 days or more650.231580.65415
First · nfl-margin-lattice-1650.231580.65415
Latest · All decided650.231190.65083
Latest · Under 24 hours630.228250.64476
Latest · 1–7 days20.323890.84222
Latest · 7 days or more0——
Latest · nfl-margin-lattice-1650.231190.65083

Missing first: 0 · missing latest: 0 · later publications in paired set: 65. 1 at or after actual kickoff.

14,498 stored snapshots · through Oct 10, 2026, 4:08 PM UTC
Results fetched through Oct 10, 2026, 4:07 PM UTC · latest result kickoff Oct 9, 2026, 12:15 AM UTC
Comparison scored Oct 10, 2026, 4:09 PM UTC · first log Oct 10, 2026, 4:09 PM UTC

Latest valid publication strictly before the stored result kickoff; old schedule timestamps do not set eligibility. Equal instants use lexical model version order, without ranking models. Conflicting probabilities for the same instant and version are withheld.

Both cohorts use stored results and full-precision conditional home-win probabilities. Kickoff and outcomes have not been independently recollected. The original first-publication record above remains unchanged.

Retained warehouse source ↗

The model is accountable

Exploratory research · not live performance

Held from production

Recent seasons, more weight

+0.00045

Brier change versus the incumbent · lower is better

95% interval -0.00017 to +0.00106

1,954 paired games · 2019–2025 · weekly block bootstrap

Held from production

Elo blend + calibration

+0.00059

Brier change versus the incumbent · lower is better

95% interval -0.00198 to +0.00314

1,136 paired games · 2022–2025 · weekly block bootstrap

Neither challenger established better probability forecasts. These experiments use the archived local corpus through the 2025 postseason; their samples differ from the production benchmark. Game probabilities remain on the incumbent model.

Better season simulations, now explorable

Actual and simulated head-to-head results now reach the playoff seeding step.

Explore playoff stakes ↗

Scored on decided games

forecasterbrierlog lossaccuracyecen
Historical market0.21200.611666.3%0.01213,871
Margin model0.22000.629464.0%0.02075,748
Elo only0.21990.629364.4%0.01285,748
Constant base rate0.24670.686556.0%0.01375,748

Historical ESPN prices have no verified closing timestamp. Ties (13) are excluded and counted — a moneyline voids on one, so every comparison is on decided games.

The margin model does not currently beat Elo alone — level on Brier, worse calibrated. Nine features have bought nothing over a rating gap and home field, and this page is not going to imply otherwise.

Paired against historical prices

Brier gap +0.00871 · 95% CI [+0.00566, +0.01170]

3,884 priced games; 1,878 unpriced and excluded rather than compared against nothing. Market probabilities come from 3,235 moneyline and 649 spread.

The market is better — the expected result for a model that carries no price information.

Calibration

Calibration is a fact about the model alone — and the property this product is actually selling.

Margin modelClosing line
perfect calibration0%0%25%25%50%50%75%75%100%100%Margin model — said 17.5%, happened 17.6% over 34 gamesMargin model — said 26.2%, happened 27.4% over 208 gamesMargin model — said 35.6%, happened 30.8% over 536 gamesMargin model — said 45.3%, happened 42.7% over 981 gamesMargin model — said 55.1%, happened 53.3% over 1,379 gamesMargin model — said 64.7%, happened 62.9% over 1,372 gamesMargin model — said 74.5%, happened 75.5% over 894 gamesMargin model — said 83.9%, happened 85.3% over 319 gamesMargin model — said 91.7%, happened 96.0% over 25 gamesClosing line — said 8.8%, happened 0.0% over 8 gamesClosing line — said 16.5%, happened 18.8% over 96 gamesClosing line — said 25.6%, happened 24.6% over 285 gamesClosing line — said 35.6%, happened 36.0% over 467 gamesClosing line — said 44.1%, happened 43.6% over 546 gamesClosing line — said 55.8%, happened 54.6% over 711 gamesClosing line — said 64.8%, happened 62.7% over 759 gamesClosing line — said 74.8%, happened 75.6% over 660 gamesClosing line — said 84.3%, happened 86.5% over 296 gamesClosing line — said 92.3%, happened 93.0% over 43 gameswhat it saidwhat happened
Dot area is the number of games in the bucket. Above the dashed line means it was too cautious; below it, too confident. The two are drawn on different game sets — the market only exists where a line was published — so read each against the diagonal rather than against the other.
Table view
ForecasterSaidHappenedGamesGap
Margin model17.5%17.6%340.2%
26.2%27.4%2081.2%
35.6%30.8%536-4.9%
45.3%42.7%981-2.6%
55.1%53.3%1,379-1.8%
64.7%62.9%1,372-1.8%
74.5%75.5%8941.0%
83.9%85.3%3191.4%
91.7%96.0%254.3%
Closing line8.8%0.0%8-8.8%
16.5%18.8%962.2%
25.6%24.6%285-1.1%
35.6%36.0%4670.4%
44.1%43.6%546-0.5%
55.8%54.6%711-1.2%
64.8%62.7%759-2.0%
74.8%75.6%6600.8%
84.3%86.5%2962.2%
92.3%93.0%430.7%

Season by season

Margin modelClosing linelower is better
0.1930.2180.243Margin model · 2005 · Brier 0.2114Margin model · 2006 · Brier 0.2352Margin model · 2007 · Brier 0.2127Margin model · 2008 · Brier 0.2195Margin model · 2009 · Brier 0.2068Margin model · 2010 · Brier 0.2324Margin model · 2011 · Brier 0.2123Margin model · 2012 · Brier 0.2215Margin model · 2013 · Brier 0.2146Margin model · 2014 · Brier 0.2091Margin model · 2015 · Brier 0.2222Margin model · 2016 · Brier 0.2181Margin model · 2017 · Brier 0.2172Margin model · 2018 · Brier 0.2180Margin model · 2019 · Brier 0.2262Margin model · 2020 · Brier 0.2194Margin model · 2021 · Brier 0.2310Margin model · 2022 · Brier 0.2242Margin model · 2023 · Brier 0.2291Margin model · 2024 · Brier 0.2142Margin model · 2025 · Brier 0.2215Margin model · 2026 · Brier 0.22900.229Closing line · 2012 · Brier 0.2279Closing line · 2013 · Brier 0.2153Closing line · 2014 · Brier 0.2045Closing line · 2015 · Brier 0.2283Closing line · 2016 · Brier 0.2151Closing line · 2017 · Brier 0.2007Closing line · 2018 · Brier 0.2119Closing line · 2019 · Brier 0.2009Closing line · 2020 · Brier 0.2011Closing line · 2021 · Brier 0.2182Closing line · 2022 · Brier 0.2089Closing line · 2023 · Brier 0.2184Closing line · 2024 · Brier 0.2008Closing line · 2025 · Brier 0.2120Closing line · 2026 · Brier 0.22900.2290508111417202326
Each point is one season scored on its own. The market line begins where the corpus first carries prices — nothing was published before about 2011, and a line drawn across that gap would invent a benchmark that did not exist.
Table view
SeasonGamesModelEloMarketGap (paired)
2026640.22900.22690.22900.0000
20252840.22150.22290.2120+0.0096
20242850.21420.21050.2008+0.0134
20232850.22910.23260.2184+0.0106
20222820.22420.22280.2089+0.0152
20212840.23100.23120.2182+0.0127
20202680.21940.21570.2011+0.0183
20192660.22620.22250.2009+0.0253
20182650.21800.22330.2119+0.0072
20172670.21720.21680.2007+0.0162
20162650.21810.21880.2151+0.0030
20152670.22220.22430.2283-0.0061
20142660.20910.20680.2045+0.0046
20132660.21460.21660.2153-0.0001
20122660.22150.21960.2279-0.0073
20112670.21230.2132——
20102670.23240.2322——
20092670.20680.2064——
20082660.21950.2184——
20072670.21270.2154——
20062670.23520.2333——
20052670.21140.2113——

The gap column is computed on the PAIRED subset — the priced games only — never by subtracting the model and market columns beside it. Those two are measured on different game sets, and in a season where the unpriced games happened to be lopsided their difference is mostly a fact about coverage.

Margin and total

Every game card publishes a margin and a total, so both are scored here.

Margin

Mean abs. error
10.53
Bias
+0.45
vs the spread
10.00
Gap
+0.32
IntervalNominalActualGap
central 50%50%54.5%+4.5
central 80%80%81.8%+1.8
central 95%95%95.1%+0.1

Errors are in points. The margin interval is read off the LATTICE, which is discrete, so it is the smallest range of whole point margins whose mass reaches the nominal level. That is conservative by construction — over- coverage at the 50% level is mostly the fat cells at 3 and 7 points, not a miscalibration.

Total

Mean abs. error
10.84
Bias
-0.81
vs the posted total
10.44
Gap
+0.40
IntervalNominalActualGap
central 50%50%51.3%+1.4
central 80%80%80.9%+0.9
central 95%95%95.1%+0.1

Errors are in points. The total is served as a normal and is measured as one.

Margin distribution shape

0%10%20%decile 1 — 10.1% of games (581), expected 10%decile 2 — 10.3% of games (591), expected 10%decile 3 — 10.1% of games (582), expected 10%decile 4 — 10.4% of games (601), expected 10%decile 5 — 10.2% of games (585), expected 10%decile 6 — 11.1% of games (639), expected 10%decile 7 — 10.7% of games (615), expected 10%decile 8 — 9.0% of games (518), expected 10%decile 9 — 8.8% of games (508), expected 10%decile 10 — 9.4% of games (542), expected 10%uniformbelow the modelmodel medianabove the model
Where the real margin fell inside the model’s own published distribution, in deciles. Flat is correct. Chi-square per degree of freedom is 3.01, and no p-value is reported — at 5,762 games any real model fails a goodness-of-fit test on some decimal place, so “p < .001” beside a visibly flat histogram would be true and completely misleading.

Total distribution shape

0%10%20%decile 1 — 7.8% of games (450), expected 10%decile 2 — 10.5% of games (604), expected 10%decile 3 — 10.6% of games (610), expected 10%decile 4 — 9.8% of games (566), expected 10%decile 5 — 11.0% of games (632), expected 10%decile 6 — 10.4% of games (602), expected 10%decile 7 — 10.1% of games (581), expected 10%decile 8 — 9.2% of games (527), expected 10%decile 9 — 9.3% of games (539), expected 10%decile 10 — 11.3% of games (651), expected 10%uniformbelow the modelmodel medianabove the model
Where the real total fell inside the model’s own published distribution, in deciles. Flat is correct. Chi-square per degree of freedom is 6.01, and no p-value is reported — at 5,762 games any real model fails a goodness-of-fit test on some decimal place, so “p < .001” beside a visibly flat histogram would be true and completely misleading.

Margin coverage and PIT are read off the LATTICE the model actually publishes, using the mid-P transform for a discrete distribution. The total is served as a normal and is measured as one. Comparing either against a normal at the same sd would grade a distribution this site never published.

A Brier here is binary, on NFL games — never comparable across the sibling projects.