Methodology
How luck is estimated, and how well the estimates hold up
We define a team’s luck in a game as the difference between its result and the probability that its own performance that day, replayed with every chance event at its expected rate, wins. That probability is estimated by Monte Carlo simulation and validated out of sample. This page sets out the estimands, the estimators and their assumptions, and every validation result, including those that did not favor the method.
The short version
Every game replayed 4,000 times from its own plays. Luck is the result minus the share of replays a team wins. Show the 3 steps and an example Hide
- Reshuffle the game and replay it. We take every snap each team ran that day and deal them out again at random into new drives, matched by down and distance: on 3rd and 7, a team gets one of its own 3rd-and-long plays from that game. A play can come up twice in one replay, or not at all. Each drive runs until a touchdown, a field goal try, a punt, a turnover or a failed 4th down, over as many possessions as the real game had. Then we see who scored more, and do it all 4,000 times.
Kept as it happened
- Every snap’s yards: runs, catches, incompletions, sacks, scrambles, and the flags on them
- How often each team fumbled, and how many interception-worthy throws its quarterback made
- The number of possessions
Drawn again
- The order the plays come in
- Whether an interception-worthy throw is caught, at the rate that defense catches them
- Who recovers a fumble (the fumbling team keeps 52%)
- Whether a field goal is good, at the league rate for its distance
No matchup is modeled on its own: there are no player ratings, no pass rush against the quarterback, no receiver against his cornerback. They don’t need to be. Every snap in the deck was played against that opponent that day, so how the pass rush fared against that quarterback, the receivers against those defenders and the line against that front is already in the plays: in the sacks, the scrambles, the completions and the stuffed runs. Where a team ran few plays of a kind that day, some draws come from its earlier games that season. Fourth downs follow one rule for everyone: kick when in range, go for it on short yardage in the other team’s half, and whenever a touchdown is needed late.
- Compare it with the result. The share of replays a team wins is what its play was worth. Luck is the result (1 for a win, 0 for a loss) minus that share. Times 100 it is a game’s luck index, from −100 to +100; added up over a season it is luck in wins.
- Find where it came from. This part is not a replay. Every chance play of the real game is scored on its own: what happened, minus how often it happens for that player in that spot, times what was at stake. Here each player’s own record counts: a kicker’s from that distance, in that wind and cold; a receiver’s drop rate; how often that defense catches interception-worthy throws; how often that offense and that defense convert on 3rd and 4th down. Judgment calls count too, at half, each weighted by how often the flagged player draws one. It doesn’t change the replay; it shows where the luck came from.
An example: Bears 24, Packers 22 Week 18 2024
Who wins the 4,000 replays
- The game
- The Packers lost 24–22, but with both teams’ plays dealt out again, they won 99% of the replays. Packers: 100 × (0 − 0.99) = −99 Bears: 100 × (1 − 0.01) = +99
- One play, scored on its own
- C. Santos made a 51-yard field goal. Going by his record, the distance and the conditions, he makes that kick 62% of the time. From that distance a make is worth 4.5 points more than a miss, on average: the 3 points, plus the field position, since after a miss the other team takes over at the spot of the kick. Luck is what happened (1 for a make, 0 for a miss) minus the 62%, times those 4.5 points: Bears: (1 − 0.62) × 4.5 = +1.7 points Counted at the moment it came, with 0:02 left in the 4th quarter, that was worth 38% of a win to the Bears.
- The season
- The Packers finished 2024 at 11–6. The replays of those 17 games add up to 12.5 wins. Luck in wins: 11 − 12.5 = −1.5
The rest of this page sets out each step in full, and how well the numbers hold up.
1. The estimand
For team i in game g, let Y ∈ {0, ½, 1} be the result (loss, tie, win) and let π be the probability that i wins a replay of g in which each team’s plays are resampled from its own performance that day and every chance event occurs at its expected, luck-neutral rate. The luck index is
L = 100 · (Y − π̂),
where π̂ is the Monte Carlo estimate of π. L lies in [−100, 100]. It is positive when a team won more than its performance warranted and negative when it won less: a loss in a game won in 92% of replays scores −92, a win at 37% scores +63. Summed over a season, Σ(Y − π̂) is luck in wins, the difference between a team’s record and the record its play would have earned on average.
The estimand is deliberately conditional on the day’s performance. It asks what that performance was worth, not whether the same teams would win again: the latter depends on everything that changes between two games (see the rematch test).
2. The replay with the same day’s form
Resampling. Each team’s offensive snaps in the game (runs, passes, and accepted penalties on them) are partitioned into nine strata by down (1st, 2nd, 3rd/4th) and distance (≤3, 4–7, 8+ yards). A replay builds drives by sequential resampling: each snap’s outcome is drawn from the stratum that matches the current down and distance, so a 3rd-and-7 is answered with one of that team’s 3rd-and-long plays from that day.
Pooling. Sparse strata are regularized by empirical-Bayes pooling. With ns same-game snaps in stratum s, a draw comes from them with probability ns / (ns + κ), κ = 12, and otherwise from the team’s earlier snaps that season, or the league’s where fewer than ten exist.
Chance at its expected rate. Turnovers occur per snap at the team’s expected rate rather than its realized one. In games with FTN charting, each interception-worthy throw counts at the opposing defense’s shrunk catch rate and every other throw at that season’s clean-throw rate; in uncharted games every throw counts at the league rate. Each fumble is lost with the league probability (48.3%). Field goals are made at the league make probability for their distance. Fourth downs follow a fixed policy: kick when in range, go for it on short yardage in opposing territory, and go for it whenever a team needs a touchdown on its last two possessions.
Monte Carlo. Each game is replayed 4,000 times with the observed number of possessions. The Monte Carlo standard error of the raw estimate is at most 0.8 percentage points, and at most 1.5 after calibration, which stretches it by up to T.
Calibration. Pooling shrinks each team toward its prior, so the raw simulator is under-dispersed: its probabilities are drawn toward ½. A single logit-scale recalibration corrects this, π̂ = σ(T · logit π̂raw), with T = 1.93 estimated by maximum likelihood on 2019–2023 and evaluated on the 2024–2025 holdout. The distribution of simulated margins on each game page is shifted so that it reproduces the calibrated probability.
3. Three reference estimates
Each game page reports three further probabilities. Each answers a different question, and each serves as a benchmark for the headline.
- Before kickoff
- The market’s ex-ante estimate: Φ(s / 13.5), with s the closing point spread.
- Same teams, fresh game
- A Bayesian measurement model. Luck-neutral efficiency (the difference in expected points per play on competitive snaps, with modeled luck removed) is treated as a noisy measurement of the teams’ latent strength that day beyond the spread. Its reliability is identified from the covariance between alternating possession pairs (split-half r = 0.21). The posterior strength is added to the spread, and one game’s residual randomness is integrated out. Latent strength on the day has a standard deviation of about 4.2 points, against 12.1 points for a single game’s randomness, so this estimate stays close to the prior. As a headline it would turn the luck index into an upset index.
- Only coin-flips re-rolled
- Conditional Monte Carlo over the enumerated chance events alone. Every kick, fumble recovery, drop and conversion is redrawn as Bernoulli(pe), and every judgment call keeps its size while a fair coin decides its direction. All other play is held fixed. It is the best season-level luck signal (validation), but too narrow for one game: it attributes all unmodeled variance to skill.
4. Event-level luck
Each chance event e is scored as the deviation of its outcome from a conditional expectation, valued in expected points (EP):
ℓe = (oe − pe) · ve,
where oe ∈ {0, 1} indicates the outcome favorable to the beneficiary, pe is its conditional probability, and ve is the mean EP difference between the two outcomes in that situation. If pe is calibrated, E[ℓe] = 0; we test this for every component each time the model is refitted. Each play contributes to at most one component: a 3rd-down pass that is dropped or intercepted, or a fumble that is lost, is scored under that component and not also as a failed conversion.
| Component | Expectation pe and value ve | Share counted as luck | Week-to-week r | Predicts later margin, r |
|---|---|---|---|---|
| Kicks | Field goals against a logistic make model (distance as a linear spline with knots at 50 and 60 yards, wind, cold), shifted by the kicker’s own record; extra points against nflfastR’s make probability, recalibrated to the league rate. | 100% | 0.15 | 0.00 |
| Fumble bounces | Which team recovers a loose ball. The fumbling team recovers 51.7%; the outcomes differ by 4.2 expected points. | 100% | 0.02 | 0.05 |
| Drops | Every catchable ball against the receiver’s own drop rate (FTN charting, 2022 on), valued by down: 1.1, 1.4, 2.5, 5.2 points on 1st to 4th down. | 100% | 0.05 | 0.02 |
| Judgment calls | Fouls that require an official’s judgment, valued at the expected points they transferred, including any play they nullified. | 50% | 0.09 | −0.17 |
| 3rd & 4th downs | 3rd- and 4th-down attempts against a logistic model of distance, down, goal-to-go and the offense’s early-down efficiency that day, shifted by both teams’ records. | 25% | −0.03 | 0.24 |
| Interceptions | Interception-worthy throws that were not intercepted and clean throws that were (FTN charting), against pick rates specific to each season’s charting. | 0% | 0.00 | 0.16 |
| Pre-snap flags | False starts, offside and similar fouls. They repeat, so they measure discipline rather than chance. | 0% | 0.01 | 0.08 |
| Special teams | Return touchdowns, blocked punts and recovered onside kicks. They predict later margins (unit quality). | 0% | 0.16 | 0.24 |
| Injuries | Absent key players, valued by position and quality, counted at three quarters of that value (section 7). Held fixed in every replay. | 75% | 0.62 | −0.01 |
| Rest | Rest-day advantage, 0.20 points per day (fitted). | 100% | −0.24 | 0.01 |
Week-to-week r: correlation between a team’s average in odd and even weeks of a season. Predicts later margin: correlation between its weeks 1–9 average and its weeks 10–18 point margin. For pure chance both should be near zero; injuries repeat because absences last for weeks.
Player- and team-specific expectations. League baselines are adjusted for the record of whoever was involved, using strictly earlier plays (no look-ahead), by empirical-Bayes shrinkage:
pe = pleague + Σj<e (oj − pleague,j) / (n + k).
The prior strength k, in pseudo-observations, is selected by out-of-sample log-likelihood on 2019–2023:
- kickers: k = 320, plus a separate long-range offset beyond 45 yards with k = 80
- receivers’ drop rates: k = 320
- defenses’ catch rates on interception-worthy throws: k = 2,560
- 3rd and 4th downs: the offense (k = 320) and the defense (k = 1,280)
A small k means an individual’s history is informative. A reliable kicker therefore carries a high pe: a make is barely lucky and a miss very unlucky. Summed over any one player’s events, luck is approximately zero.
5. Leverage: luck in win probability
Expected points measure an event’s size; win probability (WP) measures its consequence. For binary events, the realized win probability added (WPA) is the natural measure. WP is a martingale, so the realized swing of a binary event equals (o − p) times the swing between its two outcomes; we rescale it from the WP model’s implied p to ours. On 3rd and 4th down only the conversion is the chance component, so the play’s WPA is scaled to the conversion’s share of its EP, |ℓe| / |EPA|: a 70-yard touchdown on 3rd-and-2 is mostly play, not chance. Other events are converted at the local exchange rate implied by a normal model of the final margin,
∂WP/∂pt = φ(Φ−1(WP)) / (13.5 · √(t/60)),
with t the minutes remaining, capped at the play’s own |WPA| / |EPA|. Three points are worth about 9% of a win at the kickoff of an even game, most of a win in the last minute of a tied one, and nothing once the game is decided. Every value is bounded by the game state: an event cannot be worth more to a team than its win probability after the play, nor cost it more than one minus that. The bound also holds after the momentum multiplier (section 8).
The site explains luck with this value: “what the luck was worth” and the turning point on game pages, luck by component on team pages, and every referee figure. It also describes a season better. Summed over a regular season, a team’s luck in win probability correlates with its wins minus replay wins at r = 0.54, against r = 0.42 for the same events in points (224 team-seasons). The turning point is the luck play worth the most to the eventual winner between its lowest win probability and the last time its chance crossed 50%, counted at its luck share and required to be worth at least 15 points of win probability.
6. Officiating
A flag’s value is the expected points it transferred to the team it helped, including any play it nullified: a holding call that erases a 40-yard reception costs the offense the reception as well as the yardage. Each flag is weighted by the penalized player’s history, w = min(1, max(½, rposition / r̂player)), where r̂player is the player’s flag rate per snap shrunk toward his position’s (k = 3,000 snaps). A frequent offender’s flag is thus partly attributed to his own discipline, down to half its value, and no flag counts for more than it moved the game. Judgment calls count as half chance (section 7); pre-snap fouls not at all, because they repeat. In the coin-flip replay a judgment call keeps its size, and a fair coin decides which team it favors. The referee pages sum calls in win probability (section 5), so a flag in a decided game counts for nothing and one at the turn of a close game for most of a win.
7. Luck shares: how much of each component is chance
A component can mix chance and skill. The share σc of component c that behaves as chance is identified by a predictive-validity criterion: the correlation between a team’s luck-adjusted margin in weeks 1–9 and its raw margin in weeks 10–18, over 2019–2023. Removing noise improves that prediction; removing skill degrades it. Shares take values on {0, ¼, ½, ¾, 1} and are fitted by coordinate ascent from the prior σc = 1 (every modeled chance event is chance). A share moves only when the criterion rises by at least 0.003, about one bootstrap standard error over team-seasons; smaller gains are within sampling noise. The fitted shares are the same for any threshold from 0.002 to 0.005.
Kicks, fumble recoveries and drops are retained as chance. Interception luck, special teams and pre-snap fouls are estimated at 0%: they predict later margins, so they behave as skill. Third and fourth downs are mostly skill. The training fit prefers 0%, but the holdout criterion is flat between 0% and 50% and highest at 25%, which we adopt.
Two shares are set from a longer record than the five training seasons. Judgment calls count at 50%: rebuilt from play-by-play for 27 seasons (1999–2025), they come out at about a third chance, and counted in full a team’s call luck starts to repeat later in the season, which chance does not do. Injuries count at 75%: one point of injury value has moved the final margin by 0.72 ± 0.11 points (2019–2025), so a quarter of the valuation was over-counted. How the luck-adjusted record predicts the future.
On the 2024–2025 holdout, the luck-adjusted margin predicts later margins at r = 0.50, against 0.51 for the raw margin and 0.38 if everything modeled were counted as luck. The adjustment removes noise at little cost in signal.
8. Momentum
We tested one kind of momentum: whether a lucky break makes the team that got it play better afterwards, beyond what the break did to the score, field position and clock. Across 18,562 pure-chance events (fumble recoveries, kicks, judgment calls and interception luck worth at least one point), it did not. The beneficiary’s next drive was no better than average: −0.007 ± 0.016 expected points per point of luck (± one standard error). Over the rest of the game, teams won as often as the win probability right after the break said, and in matched situations in which a coin flip landed differently (a long field goal made or missed, a fumble recovered or lost), the lucky side played no better afterwards (the momentum tests).
The tests do not rule out momentum from plays a team earned, or a small effect of luck. They do rule out a large one: had each break carried half its value over, teams would have won their breaks worth 10% or more about 9 percentage points more often than the win probability said. Across 2,184 such breaks they won 0.6 ± 0.8 more.
Momentum is a setting. Every in-game luck event counts at 1 + β times its value in the coin-flip replay, in luck in points and in luck in wins, with β from none to +100%. By default β = 0, the value the tests measured. Momentum doesn’t improve the luck-adjusted record as a forecast either: coin-flip wins after nine weeks predict the rest of a season at r = 0.425 at the default and r = 0.427 at +50%, our default from 28 to 30 September 2026. Scored after every number of games from four to twelve, and game by game, +50% forecasts worse (the forecast test).
9. What the “adjust the assumptions” panels change
The panels rerun the coin-flip replay, luck in points and the luck-adjusted margin under your shares, momentum and choice of what counts. They do not move the replay with the same day’s form, because that replay already draws kicks, bounces and turnovers at luck-neutral rates, so the settings have nothing in it to act on. Flipping a play fixes it at its other outcome, in the game and in every replay.
10. Form: identifying an off day
When the replay with the same day’s form and the fresh-game estimate differ by 30 percentage points or more, the game departed sharply from what the two teams usually are. We name the team that fell short and compare each unit’s play with its usual level: expected points per play on competitive snaps, with turnovers excluded, against the team’s earlier games that season plus half of the previous season, shrunk toward the league average. One team’s defense and the other team’s offense are the same snaps, so a single game cannot fully apportion whose day it was.
11. Validation
Parameters are fitted on 2019–2023 and evaluated on 2024–2025 unless noted; 1,960 games are simulated. Source: python -m engine.sports.nfl.research validate, run 2026-09-30.
When the replay says 90%, it happens about 90% of the time
Every game 2019–2025, grouped by the replay’s win probability for the home team: what it predicted against how often the home team actually won.
Show as a table
| Bin | Games | Predicted | Actually won |
|---|---|---|---|
| 0%–10% | 223 | 4.8% | 5.4% |
| 10%–25% | 290 | 17.0% | 18.3% |
| 25%–50% | 389 | 37.6% | 38.8% |
| 50%–75% | 404 | 62.6% | 61.0% |
| 75%–90% | 340 | 82.9% | 86.2% |
| 90%–100% | 314 | 95.3% | 94.4% |
Does it predict the rest of a season?
Correlation between a team’s weeks 1–9 value and its win share in weeks 10–18. The coin-flip replay predicts best, above the closing line; the replay with the same day’s form does no better than actual wins. Season forecasts on this site therefore use the coin-flip number.
| Coin-flip wins (at the default momentum) | 0.425 |
| Point differential | 0.400 |
| Post-game win expectancy (success rate and EPA) | 0.396 |
| Closing line | 0.395 |
| Actual wins | 0.378 |
| Replay wins (same day’s form) | 0.375 |
Does it predict a real rematch?
A natural experiment: 336 same-season division rematches. Does game 1 predict the winner of game 2? Leave-one-season-out log loss, lower is better; a coin flip scores 0.693. Nothing from one game predicts the next much better than a coin flip. The replay matches the closing line, within the sampling noise of this many pairs, which is consistent with the estimand: a replay claim means “same day, same performances”, not “they would win next time”.
| Replay with same day’s form | 0.666 |
| Closing line | 0.668 |
| Game 1 margin | 0.680 |
| Coin-flip replay | 0.682 |
| Game 1 result | 0.687 |
Market efficiency
Season-to-date luck does not predict covering the next spread (r = 0.001, 3,248 games). The closing line already prices this kind of luck (luck and the spread).
Unbiasedness
Each time the model is refitted, every binary component is tested for E[ℓe] = 0 by comparing the realized frequency of the favorable outcome with the mean expectation. Every component passes at the 5% level (|z| ≤ 1.4), including field goals of 60 yards or more, except interception luck, whose stored events are selected by design (limitations).
12. Corrections
28 September 2026: an audit of the luck engine. A review of the model found the following, each now corrected, refitted and revalidated. The numbers on this page are after the corrections.
- Long field goals. The make model was quadratic in distance, which flattens past 60 yards: it expected 53% on kicks of 60+ yards, which were made 39% of the time, so long misses were over-counted as bad luck. Distance is now a monotone linear spline, and those kicks are calibrated.
- Expected interceptions in the replay. In charted games, a team that made no interception-worthy throw was assigned the league interception rate (about 0.8 expected interceptions instead of about 0.1), and clean throws were left out of everyone else’s expectation. FTN’s charting standard also tightened over time, so interception rates are now estimated for each charting season.
- Conversions in win probability. A conversion’s luck was valued at the whole play’s swing, including every yard past the sticks. It is now limited to the conversion’s share of the play.
- Double counting. A dropped or intercepted 3rd-down pass, or a lost fumble, was scored twice: under its own component and again as a failed conversion. Each play now counts once, and drops are valued by down (a 3rd-down drop costs a conversion; the single value used before over-valued 1st-down drops by half).
- Flag weights. A flag on a rarely penalized player counted up to twice its value, more than the call moved the game, and broke the win-probability bound. The weight is now at most 1.
- Bounds under momentum. The momentum multiplier could push a single event past 100% of a win. Values are now bounded by the game state after the multiplier.
- Calibration of two baselines. nflfastR’s extra-point probability ran about 1.3 points low, so every made extra point was booked as slightly lucky; it is now recalibrated. Fumble recoveries now use the measured rate (51.7% to the fumbling team), not exactly half.
- Share fitting. The share search accepted any improvement, however small. After the corrections above, noise alone would have set drops to 0% luck, so a share now moves only for a gain above sampling noise.
September 2026: team codes. The 2019 Raiders were filed under two team codes in different sources, so their luck was mis-signed and their games had no replay number. Fixed, refitted and revalidated.
13. Data
Play-by-play, win probability, schedules and closing lines, snap counts, rosters and injury reports come from nflverse (CC BY 4.0), updated each morning in season. Drops and interception-worthy throws come from FTN Data via nflverse (CC BY-SA 4.0), from 2022; earlier games are replayed without them, and their pages say so. This site’s data is shared under CC BY-SA 4.0.
14. Limitations
- Late-game win probabilities come from nflfastR’s model, which is least reliable in the final seconds. A flag that nullifies a late play is valued as the call plus the erased play, which can overstate its swing; displayed swings are capped at 100 points.
- Judgment calls count at a fixed share rather than against an expected flag rate given each team’s style of play. Teams that throw deep draw more pass interference, for example. Their low week-to-week repeatability (0.09) suggests this matters little over a season.
- Only notable throws are stored as interception-luck events: interception-worthy throws, and clean throws that were picked. Those events are therefore negative on average. Interception luck is counted at 0%, so this affects nothing at the default settings; the replay’s turnover rates account for every throw.
- In the season under way, interception rates use the latest complete season’s charting until the season is refitted.
- Player quality for injuries comes from Approximate Value in the draft data, a coarse measure that includes seasons after the game.
- Luck shares are identified from about 160 training team-seasons and are coarse by design (quarters). Interception luck has only two training seasons of charting.
- A game appears the morning after it is played. FTN charting can take 48 hours, and the page updates when it arrives.