Held from production
Recent seasons, more weight
+0.00045
Brier change versus the incumbent · lower is better
95% interval -0.00017 to +0.00106
1,954 paired games · 2019–2025 · weekly block bootstrap
walk-forward, weekly, expanding window
Historical backtest5,762 games scored from 2005 · 3-season warm-up · de-vig shin · published Oct 6, 2026
Earliest recorded pregame forecast per fixture. This small live sample is separate from the historical backtest below. Published Oct 10, 2026.
model version
| Cohort | Games | Brier | Log loss |
|---|---|---|---|
| nfl-margin-lattice-1 | 65 | 0.2316 | 0.6542 |
horizon
| Cohort | Games | Brier | Log loss |
|---|---|---|---|
| under 24h | 0 | — | — |
| 1 to 7 days | 0 | — | — |
| 7 days or more | 65 | 0.2316 | 0.6542 |
Retained publications · 2026
Small sample65 identical decided games · probability scores, lower is better.
Latest − first Brier: -0.00039 · descriptive difference only.
2 decided games have no valid stored forecast under 24 hours before kickoff. This sample does not establish an accuracy gain or support model promotion.
All-cohort scores below can have different game IDs. Only the paired cards above compare identical decided games; ties are counted and excluded.
| Publication / cohort | Games | Brier | Log loss |
|---|---|---|---|
| First · All decided | 65 | 0.23158 | 0.65415 |
| First · Under 24 hours | 0 | — | — |
| First · 1–7 days | 0 | — | — |
| First · 7 days or more | 65 | 0.23158 | 0.65415 |
| First · nfl-margin-lattice-1 | 65 | 0.23158 | 0.65415 |
| Latest · All decided | 65 | 0.23119 | 0.65083 |
| Latest · Under 24 hours | 63 | 0.22825 | 0.64476 |
| Latest · 1–7 days | 2 | 0.32389 | 0.84222 |
| Latest · 7 days or more | 0 | — | — |
| Latest · nfl-margin-lattice-1 | 65 | 0.23119 | 0.65083 |
Missing first: 0 · missing latest: 0 · later publications in paired set: 65. 1 at or after actual kickoff.
14,498 stored snapshots · through Oct 10, 2026, 4:08 PM UTC
Results fetched through Oct 10, 2026, 4:07 PM UTC · latest result kickoff Oct 9, 2026, 12:15 AM UTC
Comparison scored Oct 10, 2026, 4:09 PM UTC · first log Oct 10, 2026, 4:09 PM UTC
Latest valid publication strictly before the stored result kickoff; old schedule timestamps do not set eligibility. Equal instants use lexical model version order, without ranking models. Conflicting probabilities for the same instant and version are withheld.
Both cohorts use stored results and full-precision conditional home-win probabilities. Kickoff and outcomes have not been independently recollected. The original first-publication record above remains unchanged.
Retained warehouse source ↗Held from production
+0.00045
Brier change versus the incumbent · lower is better
95% interval -0.00017 to +0.00106
1,954 paired games · 2019–2025 · weekly block bootstrap
Held from production
+0.00059
Brier change versus the incumbent · lower is better
95% interval -0.00198 to +0.00314
1,136 paired games · 2022–2025 · weekly block bootstrap
Neither challenger established better probability forecasts. These experiments use the archived local corpus through the 2025 postseason; their samples differ from the production benchmark. Game probabilities remain on the incumbent model.
Better season simulations, now explorable
Actual and simulated head-to-head results now reach the playoff seeding step.
| forecaster | brier | log loss | accuracy | ece | n |
|---|---|---|---|---|---|
| Historical market | 0.2120 | 0.6116 | 66.3% | 0.0121 | 3,871 |
| Margin model | 0.2200 | 0.6294 | 64.0% | 0.0207 | 5,748 |
| Elo only | 0.2199 | 0.6293 | 64.4% | 0.0128 | 5,748 |
| Constant base rate | 0.2467 | 0.6865 | 56.0% | 0.0137 | 5,748 |
Historical ESPN prices have no verified closing timestamp. Ties (13) are excluded and counted — a moneyline voids on one, so every comparison is on decided games.
The margin model does not currently beat Elo alone — level on Brier, worse calibrated. Nine features have bought nothing over a rating gap and home field, and this page is not going to imply otherwise.
Brier gap +0.00871 · 95% CI [+0.00566, +0.01170]
3,884 priced games; 1,878 unpriced and excluded rather than compared against nothing. Market probabilities come from 3,235 moneyline and 649 spread.
The market is better — the expected result for a model that carries no price information.
Calibration is a fact about the model alone — and the property this product is actually selling.
| Forecaster | Said | Happened | Games | Gap |
|---|---|---|---|---|
| Margin model | 17.5% | 17.6% | 34 | 0.2% |
| 26.2% | 27.4% | 208 | 1.2% | |
| 35.6% | 30.8% | 536 | -4.9% | |
| 45.3% | 42.7% | 981 | -2.6% | |
| 55.1% | 53.3% | 1,379 | -1.8% | |
| 64.7% | 62.9% | 1,372 | -1.8% | |
| 74.5% | 75.5% | 894 | 1.0% | |
| 83.9% | 85.3% | 319 | 1.4% | |
| 91.7% | 96.0% | 25 | 4.3% | |
| Closing line | 8.8% | 0.0% | 8 | -8.8% |
| 16.5% | 18.8% | 96 | 2.2% | |
| 25.6% | 24.6% | 285 | -1.1% | |
| 35.6% | 36.0% | 467 | 0.4% | |
| 44.1% | 43.6% | 546 | -0.5% | |
| 55.8% | 54.6% | 711 | -1.2% | |
| 64.8% | 62.7% | 759 | -2.0% | |
| 74.8% | 75.6% | 660 | 0.8% | |
| 84.3% | 86.5% | 296 | 2.2% | |
| 92.3% | 93.0% | 43 | 0.7% |
| Season | Games | Model | Elo | Market | Gap (paired) |
|---|---|---|---|---|---|
| 2026 | 64 | 0.2290 | 0.2269 | 0.2290 | 0.0000 |
| 2025 | 284 | 0.2215 | 0.2229 | 0.2120 | +0.0096 |
| 2024 | 285 | 0.2142 | 0.2105 | 0.2008 | +0.0134 |
| 2023 | 285 | 0.2291 | 0.2326 | 0.2184 | +0.0106 |
| 2022 | 282 | 0.2242 | 0.2228 | 0.2089 | +0.0152 |
| 2021 | 284 | 0.2310 | 0.2312 | 0.2182 | +0.0127 |
| 2020 | 268 | 0.2194 | 0.2157 | 0.2011 | +0.0183 |
| 2019 | 266 | 0.2262 | 0.2225 | 0.2009 | +0.0253 |
| 2018 | 265 | 0.2180 | 0.2233 | 0.2119 | +0.0072 |
| 2017 | 267 | 0.2172 | 0.2168 | 0.2007 | +0.0162 |
| 2016 | 265 | 0.2181 | 0.2188 | 0.2151 | +0.0030 |
| 2015 | 267 | 0.2222 | 0.2243 | 0.2283 | -0.0061 |
| 2014 | 266 | 0.2091 | 0.2068 | 0.2045 | +0.0046 |
| 2013 | 266 | 0.2146 | 0.2166 | 0.2153 | -0.0001 |
| 2012 | 266 | 0.2215 | 0.2196 | 0.2279 | -0.0073 |
| 2011 | 267 | 0.2123 | 0.2132 | — | — |
| 2010 | 267 | 0.2324 | 0.2322 | — | — |
| 2009 | 267 | 0.2068 | 0.2064 | — | — |
| 2008 | 266 | 0.2195 | 0.2184 | — | — |
| 2007 | 267 | 0.2127 | 0.2154 | — | — |
| 2006 | 267 | 0.2352 | 0.2333 | — | — |
| 2005 | 267 | 0.2114 | 0.2113 | — | — |
The gap column is computed on the PAIRED subset — the priced games only — never by subtracting the model and market columns beside it. Those two are measured on different game sets, and in a season where the unpriced games happened to be lopsided their difference is mostly a fact about coverage.
Every game card publishes a margin and a total, so both are scored here.
| Interval | Nominal | Actual | Gap |
|---|---|---|---|
| central 50% | 50% | 54.5% | +4.5 |
| central 80% | 80% | 81.8% | +1.8 |
| central 95% | 95% | 95.1% | +0.1 |
Errors are in points. The margin interval is read off the LATTICE, which is discrete, so it is the smallest range of whole point margins whose mass reaches the nominal level. That is conservative by construction — over- coverage at the 50% level is mostly the fat cells at 3 and 7 points, not a miscalibration.
| Interval | Nominal | Actual | Gap |
|---|---|---|---|
| central 50% | 50% | 51.3% | +1.4 |
| central 80% | 80% | 80.9% | +0.9 |
| central 95% | 95% | 95.1% | +0.1 |
Errors are in points. The total is served as a normal and is measured as one.
Margin coverage and PIT are read off the LATTICE the model actually publishes, using the mid-P transform for a discrete distribution. The total is served as a normal and is measured as one. Comparing either against a normal at the same sd would grade a distribution this site never published.
A Brier here is binary, on NFL games — never comparable across the sibling projects.