luckadjusted

Methodology

How the replays work, and how well they hold up

Every game is replayed 4,000 times from its own plays. How often a team wins those replays is what its performance was worth that day; the difference from the result is luck. Below: how it’s done, the three other numbers we show next to it, and every test we ran, including the ones that disappointed.

The replay with the same day’s form

For each game, each team’s own offensive plays, and the penalties on them, go into a pool sorted by down and distance. A replay draws from those pools to build new drives: a 3rd-and-7 is answered with one of that team’s 3rd-and-long plays from that day. Thin pools are topped up: with n plays of that day in a situation, a draw comes from them with weight n/(n+12), otherwise from the team’s earlier plays that season (the league’s, early in the year). Field goals use the league’s make curve, turnovers happen at the rate the team earned (expected interceptions from its interception-worthy throws, half of its fumbles) rather than the rate that happened, and trailing teams go for it late. Each game is replayed 4,000 times with as many possessions as the real one.

A single scale (1.87 on the log-odds), fitted on 2019–2023 and checked on 2024–2025, corrects the replay’s tendency to hedge toward 50%. The chart of simulated margins on each game page is shifted by the same correction so it agrees with the headline number.

The luck index is 100 × (result − replay win probability). A loss in a game a team wins 92% of the time in replays scores −92; a win at 37% scores +63.

Three other numbers, three other questions

Before kickoff
The closing betting spread, turned into a probability (a normal spread of 13.5 points). What everyone knew before the game.
Same teams, fresh game
The spread plus what this game showed about the teams, shrunk by how much one game can tell. We measured that by splitting every game into alternating possessions: beyond the spread, a team’s strength on the day varies by about ±4.3 points, while one game’s randomness is ±12.1. So this number stays close to the line, and it can’t be the headline: it would turn the luck index into an upset index.
Only coin-flips re-rolled
Every kick, fumble bounce, drop, conversion and judgment call redrawn from its own odds, everything else kept. The best measure of season-level luck (see validation), but too narrow for one game: it treats everything it doesn’t redraw as earned.

Luck, play by play

For every event, luck is what happened minus what was expected, in expected points. Expectations are specific to the players involved, from their earlier plays only: a kicker who rarely misses from 50 is expected to make it, so a miss is very unlucky and a make barely lucky.

KindWhat’s expectedCounted as luckRepeats week to week
KicksEach field goal against that kicker’s own make rate at that distance, with wind and cold; extra points against the league rate.100%0.15
Fumble bouncesWho recovers a loose ball is close to a coin flip; recoveries are valued at half a turnover either way.75%0.03
DropsEvery catchable ball against the receiver’s own drop rate (FTN charting, 2022+).100%0.05
Judgment callsFlags an official has to judge, valued in points for the team they helped, including any play they wiped out.100%0.06
3rd & 4th downs3rd and 4th downs against the chance of converting given distance, field position and how well that offense moved the ball that day.25%-0.05
InterceptionsInterception-worthy throws that weren’t picked, and clean throws that were (FTN charting).0%0.01
Pre-snap flagsFalse starts, offside and the like: mostly discipline.0%0.01
Special teamsReturn touchdowns, blocked punts, recovered onside kicks.0%0.16
InjuriesKey players missing, valued by position and quality. The same players miss every replay.100%0.62
RestDays of rest advantage, 0.2 points a day (fitted).100%-0.24

“Repeats week to week”: the correlation of a team’s average in odd and even weeks. Pure luck should repeat at about zero; injuries repeat because they last for weeks.

How referee calls are valued

Each flag is worth the expected points it gave the team it helped, including the play it wiped out: a holding call that erases a 40-yard catch costs the offense the catch as well as the yards. Flags on players who are rarely flagged count more (up to twice), flags on frequent offenders less. Judgment calls are counted fully as luck; pre-snap flags aren’t, because they track discipline.

How much of each kind of play is luck

Only part of each component is luck. The shares above were fitted on 2019–2023: the share that makes a team’s luck-adjusted margin in weeks 1–9 best predict its margin in weeks 10–18. Removing noise helps that prediction; removing skill hurts it. Kicks, drops and judgment calls come out as luck. Third downs are mostly skill (25% luck). Interception luck, special teams and pre-snap flags come out at 0%: they predict a team’s later margins, so they behave like skill.

What the “adjust the assumptions” panels change

The panels rerun the coin-flip replay, the luck in points and the luck-adjusted margin under your settings. They don’t move the replay with the same day’s form, because that replay already draws kicks, bounces and turnovers at luck-neutral rates: there’s nothing in it for the settings to change. Flipping a play fixes it at the other outcome in the game and in every replay.

Form: an off day

When the replay with the same day’s form and the fresh-game number are 30 or more points apart, the game went very differently from what the two teams usually are. We name the team that fell short and show each unit’s play in that game (expected points per play on competitive snaps, turnovers left out) against its usual level: its earlier games that season plus half of the last, pulled toward the league average. One team’s defense and the other team’s offense are the same snaps, so one game can’t fully say whose day it was.

Validation

Fitted on 2019–2023 and checked on 2024–2025 unless noted; 1,960 games. Source: python -m engine.sports.nfl.research research and sim2, run 2026-09-25 after the team-code fix (docs/METHOD.md).

When the replay says 90%, it happens about 90% of the time

Every game 2019–2025, grouped by the replay’s win probability for the home team: what it predicted against how often the home team actually won.

Show as a table
BinGamesPredictedActually won
0%–10%2075.2%5.8%
10%–25%28917.4%21.1%
25%–50%39437.6%37.6%
50%–75%43762.3%60.5%
75%–90%35983.3%86.1%
90%–100%27495.1%94.0%

Does it predict the rest of a season?

Correlation between a team’s weeks 1–9 numbers and its win share in weeks 10–18, over 224 team-seasons. The coin-flip replay predicts best; the replay with the same day’s form does no better than actual wins, which is why season forecasts on this site use the coin-flip number.

Coin-flip wins0.437
Point differential0.400
Post-game win expectancy (success rate and EPA)0.396
Betting line0.395
Actual wins0.378
Replay wins (same day’s form)0.374

Does it predict a real rematch?

336 division rematches: does game 1 predict who wins game 2? Log loss, lower is better; a coin flip scores 0.693. Nothing from one game predicts the next much better than a coin flip. The replay comes close to the betting line without beating it, which is the honest reading: a replay claim means “same day, same performances”, not “they’d win next time”.

Betting line0.668
Replay with same day’s form0.670
Game 1 margin0.680
Coin-flip replay0.682
Game 1 result0.687

Momentum and the betting market

A lucky break doesn’t carry over: across 19,080 pure-luck events the lucky team’s next drive was no better than average (momentum factor -0.005 ± 0.016 per point); the momentum test. And season-to-date luck doesn’t predict covering the next spread (correlation -0.009, 3,248 games): the market already prices it; luck and the spread.

Corrections

September 2026: the 2019 Raiders were filed under two team codes in different sources, so their luck was mis-signed and their games had no replay number. Fixed, refitted and revalidated; the numbers on this page are after the fix.

Data

Play-by-play, win probability, schedules and closing lines, snap counts, rosters and injury reports from nflverse (CC BY 4.0), updated each morning in season. Drops and interception-worthy throws: FTN Data via nflverse (CC BY-SA 4.0), from 2022; earlier games are replayed without them and say so. This site’s data is shared under CC BY-SA 4.0.

Limits

  • Late-game win probabilities come from nflfastR’s model, which is shakiest in the final seconds. A flag that wipes out a late play is valued as the call plus the erased play, which can overstate its swing; displayed swings are capped at 100 points.
  • Player quality for injuries comes from draft-data Approximate Value, a rough measure.
  • A game appears the morning after it’s played; FTN charting can take 48 hours, and the page updates when it arrives.
  • A replay can’t say whether a team would win the next meeting: that depends on everything that changes in between.