Learn · deserved to win
What is post-game win expectancy?
Post-game win expectancy asks how often a team would win a game, given how it played. Bill Connelly built it for college football from a game’s key stats, and FTN publishes an NFL version. We answer the same question by replaying each game from its own plays, 4,000 times.
How luck is calculated, in short
Every game replayed 4,000 times from its own plays. Luck is the result minus the share of replays a team wins. Show the 3 steps and an example Hide
- Reshuffle the game and replay it. We take every snap each team ran that day and deal them out again at random into new drives, matched by down and distance: on 3rd and 7, a team gets one of its own 3rd-and-long plays from that game. A play can come up twice in one replay, or not at all. Each drive runs until a touchdown, a field goal try, a punt, a turnover or a failed 4th down, over as many possessions as the real game had. Then we see who scored more, and do it all 4,000 times.
Kept as it happened
- Every snap’s yards: runs, catches, incompletions, sacks, scrambles, and the flags on them
- How often each team fumbled, and how many interception-worthy throws its quarterback made
- The number of possessions
Drawn again
- The order the plays come in
- Whether an interception-worthy throw is caught, at the rate that defense catches them
- Who recovers a fumble (the fumbling team keeps 52%)
- Whether a field goal is good, at the league rate for its distance
No matchup is modeled on its own: there are no player ratings, no pass rush against the quarterback, no receiver against his cornerback. They don’t need to be. Every snap in the deck was played against that opponent that day, so how the pass rush fared against that quarterback, the receivers against those defenders and the line against that front is already in the plays: in the sacks, the scrambles, the completions and the stuffed runs. Where a team ran few plays of a kind that day, some draws come from its earlier games that season. Fourth downs follow one rule for everyone: kick when in range, go for it on short yardage in the other team’s half, and whenever a touchdown is needed late.
- Compare it with the result. The share of replays a team wins is what its play was worth. Luck is the result (1 for a win, 0 for a loss) minus that share. Times 100 it is a game’s luck index, from −100 to +100; added up over a season it is luck in wins.
- Find where it came from. This part is not a replay. Every chance play of the real game is scored on its own: what happened, minus how often it happens for that player in that spot, times what was at stake. Here each player’s own record counts: a kicker’s from that distance, in that wind and cold; a receiver’s drop rate; how often that defense catches interception-worthy throws; how often that offense and that defense convert on 3rd and 4th down. Judgment calls count too, at half, each weighted by how often the flagged player draws one. It doesn’t change the replay; it shows where the luck came from.
An example: Bears 24, Packers 22 Week 18 2024
Who wins the 4,000 replays
- The game
- The Packers lost 24–22, but with both teams’ plays dealt out again, they won 99% of the replays. Packers: 100 × (0 − 0.99) = −99 Bears: 100 × (1 − 0.01) = +99
- One play, scored on its own
- C. Santos made a 51-yard field goal. Going by his record, the distance and the conditions, he makes that kick 62% of the time. From that distance a make is worth 4.5 points more than a miss, on average: the 3 points, plus the field position, since after a miss the other team takes over at the spot of the kick. Luck is what happened (1 for a make, 0 for a miss) minus the 62%, times those 4.5 points: Bears: (1 − 0.62) × 4.5 = +1.7 points Counted at the moment it came, with 0:02 left in the 4th quarter, that was worth 38% of a win to the Bears.
- The season
- The Packers finished 2024 at 11–6. The replays of those 17 games add up to 12.5 wins. Luck in wins: 11 − 12.5 = −1.5
Each step in full, and how well the numbers hold up: methodology.
The classic version
Take the stats of one game (success rate, yards per play, expected points added per play, turnovers) and fit a model of who won. Plug in a game, and the model says how often a team with those numbers wins. If a team lost a game its numbers usually win, it was unlucky.
Why it looks better than it is
Those stats contain the result. Expected points added counts touchdowns as big plays, so a team that scored more almost always “played better” by that measure. A model built on them predicts the winner of the same game very well (log loss 0.287 on our 2024–2025 holdout, where lower is closer) because it partly re-describes the score. Asked a question it can’t see the answer to, whether its numbers predict the rest of a season, it does barely better than actual wins (0.396 against 0.378).
The replay instead
We keep each team’s plays and throw away the order they came in. The plays are drawn again, by down and distance, into new drives and new games, with kicks and turnovers at the rates the teams earned rather than the ones that happened. The share of replays a team wins is what its play was worth that day, and it is calibrated: when it says 90%, the team wins about 90% of the time (chart).
The biggest heist since 2019 by that measure: the Bears beat the Packers 24–22 (2024, Week 18) in a game they win 1% of the time.