Skip to content
Gridiron

Method

How it works

A margin-and-total model over every NFL game since the league realigned to 32 teams in 2002, refit weekly, and scored against the closing line on named games. It is behind the market by a published margin, and that is the result the design predicts rather than a disappointment.

Games scored

5,762

walk-forward, out of sample

Gap to the close

+0.0087

Brier, lower is better

Calibration error

0.0207

expected vs observed

Priced games

3,884

paired against the close

Method

01What it does

Four things, and nothing else. If a proposed feature is none of them, it does not belong here.

  1. A win probability and expected score for every game.
  2. A projected season: record, division, seed and Super Bowl odds.
  3. A value surface comparing the model against the no-vig market price.
  4. The playoff picture — who makes the field and who hosts.

02Football margins are lumpy

Every summary statistic says a normal distribution fits these margins. Every one of them is missing the point.

Margin skewness is +0.07 and excess kurtosis is +0.20 — textbook “normal is an excellent fit”. And the distribution is nothing of the sort, because football scores are built out of 3s and 7s.

MarginActualA normal expects
314.82%5.4%
79.08%5.2%
53.57%5.4%
91.54%4.9%

A three-point game is nine times more likely than a nine-point game. No moment up to fourth order can see that, because the lumpiness is periodic rather than skewed or heavy-tailed: it moves mass between adjacent integers while leaving the shape at scale untouched.

So the model is a normal kernel modulated by a measured lattice weight. The normal carries location and spread, which do move with team strength; the weight carries the arithmetic of scoring, which does not — a three-point game is over-represented whether the game was a pick'em or a mismatch.

The weight is measured, not designed: w(3) = 2.88, w(7) = 1.83, w(9) = 0.46. The one to read twice is w(0) = 0.13. A fitted normal expects 168 ties in this corpus. There were 15.

You can see the whole lattice on any game page.

03Why that matters at the number

When the line is exactly −3, a bet can win, lose or push — and the push is worth roughly one game in twelve.

Any continuous model assigns a push zero probability by construction, then silently redistributes that mass to the two sides. On the most heavily traded number in the sport. This model prices all three outcomes, and a half-point line correctly pushes with probability zero because no integer margin equals it.

04Ties are real, and the market does not price them

An NFL game can end level. It is rare — 0.241% of regular-season games — and it is structural rather than a rounding error.

A moneyline voids on a tie: stakes returned, no winner. So a de-vigged two-way price is not P(home wins), it is P(home wins given the game is decided). Comparing an unconditional model probability against it understates the model by exactly the tie mass on every single game — a small, one-directional bias that would look like systematic shading toward the underdog.

Model probabilities are conditioned before they meet a price, and ties are excluded from every scored figure and counted beside it. The sibling NBA project has no tie branch at all and is right not to: it measures zero ties in 27,690 games.

05Home advantage has halved

A fixed constant would mis-price the modern game badly. It has fallen from about 53 rating points to about 28.

EraHome win rateMean marginRating points
2002–2006.5758+2.5653
2007–2011.5680+2.5648
2012–2016.5705+2.4351
2017–2020.5460+1.2133
2021–2025.5382+2.0828

The 2017–2020 row contains the empty-stadium 2020 season and its +1.21 should not be read as a trend point. The served model refits weekly and picks the current value up through its intercept, so it tracks this drift rather than assuming it away.

06Seeding is not by record

Four division winners take seeds 1–4 whatever the records say. A 9-8 division champion hosts a 13-4 wild card.

This is why the projection cannot be a conference table sorted by wins, which is what the sibling basketball project correctly does for its own sport. The field also changed size in 2020 — from 12 teams with two byes per conference to 14 with one — and the bracket reseeds after every round, so who a team meets in the divisional round depends on games it was not playing in.

The Super Bowl is simulated at a neutral site. Every other postseason game is hosted by the better seed, and carrying home advantage into the last one would hand the higher-rated conference champion a few percent it has not earned.

07Season simulation

Each simulated season draws one strength offset per franchise and holds it for all seventeen games, rather than perturbing each game independently.

A team that is better than its rating is better in all of them, so no number of simulations averages that error away. The offset's size is measured: within-season rating drift over 768 team-seasons has a standard deviation of 34.5 rating points.

Projected win totals are deliberately wide. Over seventeen games even a perfectly known .600 team has a binomial standard deviation of about two wins — more than a fifth of its expected total. A narrow interval here would be wrong rather than confident.

Evidence

08The market is the benchmark

The closing line beats this model, significantly, and that is the wanted result.

ForecasterBrierAccuracyECE
Market (closing line)0.212066.3%0.0121
Elo only0.219964.4%0.0128
Margin model0.220064.0%0.0207
Constant base rate0.246756.0%0.0137

The model carries no market features. A forecaster that had never seen a price and beat the price would be a bug announcing itself, not an edge — so the benchmark script logs a warning rather than a triumph if it ever happens.

Paired bootstrap on 3,884 priced, decided games: +0.00871, 95% CI [+0.00566, +0.01170]. The model closes about three quarters of the distance from a constant base rate to the market.

09The extra features have not earned their place

The nine-feature model does not beat plain Elo. That is reported rather than dressed up.

It is level on Brier and materially worse calibrated. Elo already encodes team strength, and rest, division and recent form turn out to be either small or already priced into the rating. A baseline that a model cannot beat stays live as the yardstick — baselines are never deleted here.

10A backtest is never a live record

Everything on the accuracy page is a reconstruction. The live record is a separate number and starts at zero.

The walk-forward refits on games strictly earlier than each week it scores, so the model never saw the game it is graded on. But nobody read those numbers before those kickoffs either. Every published forecast is stamped before its kickoff and graded separately, and the two are never merged however tempting a larger sample is.

Limits

11What it will not do

It will not tell you what to bet
These are model probabilities published for their own sake. The market is better than this model and the accuracy page says so with a confidence interval.
It will not show confidence it has not measured
Displayed confidence never exceeds measured confidence.
It will not fill a gap with a plausible number
Sparse coverage stays genuinely missing. “No line published” and “no edge” are different facts and render differently.

12What is missing

No injury or roster data
The model does not know who is playing. In a sport where one position is worth several points a game this is the largest single gap, and it is why preseason Super Bowl odds here stay more concentrated than a real futures market. Game pages show the injury report from ESPN so a reader can apply what the model cannot.
No team box-score features
Turnover margin and yardage were built and then removed, because the columns behind them are empty for the whole corpus. A constant feature is not a weak feature, it is an absent one.
Tiebreakers are approximated
Win percentage, head-to-head, division and conference record are modelled. The league's procedure has twelve steps; the remainder breaks deterministically rather than by simulated coin toss, because a random tiebreak inside a Monte Carlo adds variance that looks like uncertainty and is not.
The market benchmark is thin before 2012
ESPN kept no odds for the early seasons, and a backfilled line is not a closing line — it arrives with no timestamp saying when it was current.

model nfl-margin-lattice-1 · every figure on this page is read from a published artifact, not typed in