luckadjusted
NFL Week 4 · 2026

Learn · proof

Is our luck model right? Ask the future

Yes, and the test is the future. If what we call luck really is luck, it can’t carry over, so a team’s record adjusted for luck has to predict its next games better than the standings do. It does: out of every 100 NFL games, the standings pick about 62 winners right, the luck-adjusted record 63, and our full forecast, which adds starting quarterbacks, injury reports and last season, 65.

Three more in 100 sounds like a small step. In the NFL it is a big one, because so much of every game is luck: in 22 of every 100 games the team that played worse still won, and a team that is a full touchdown better still loses about 3 times in 10. Even a forecast that knew exactly how good both teams were would pick only about 68 of 100, and guessing gets 50. The standings already get to 62, so only about 6 winners in 100 are left for any forecast to find, and ours finds more than half of them: about 9 more winners called right over a 272-game season.

How luck is calculated, in short

Every game replayed 4,000 times from its own plays. Luck is the result minus the share of replays a team wins. Show the 3 steps and an example Hide
  1. Reshuffle the game and replay it. We take every snap each team ran that day and deal them out again at random into new drives, matched by down and distance: on 3rd and 7, a team gets one of its own 3rd-and-long plays from that game. A play can come up twice in one replay, or not at all. Each drive runs until a touchdown, a field goal try, a punt, a turnover or a failed 4th down, over as many possessions as the real game had. Then we see who scored more, and do it all 4,000 times.

    Kept as it happened

    • Every snap’s yards: runs, catches, incompletions, sacks, scrambles, and the flags on them
    • How often each team fumbled, and how many interception-worthy throws its quarterback made
    • The number of possessions

    Drawn again

    • The order the plays come in
    • Whether an interception-worthy throw is caught, at the rate that defense catches them
    • Who recovers a fumble (the fumbling team keeps 52%)
    • Whether a field goal is good, at the league rate for its distance

    No matchup is modeled on its own: there are no player ratings, no pass rush against the quarterback, no receiver against his cornerback. They don’t need to be. Every snap in the deck was played against that opponent that day, so how the pass rush fared against that quarterback, the receivers against those defenders and the line against that front is already in the plays: in the sacks, the scrambles, the completions and the stuffed runs. Where a team ran few plays of a kind that day, some draws come from its earlier games that season. Fourth downs follow one rule for everyone: kick when in range, go for it on short yardage in the other team’s half, and whenever a touchdown is needed late.

  2. Compare it with the result. The share of replays a team wins is what its play was worth. Luck is the result (1 for a win, 0 for a loss) minus that share. Times 100 it is a game’s luck index, from −100 to +100; added up over a season it is luck in wins.
  3. Find where it came from. This part is not a replay. Every chance play of the real game is scored on its own: what happened, minus how often it happens for that player in that spot, times what was at stake. Here each player’s own record counts: a kicker’s from that distance, in that wind and cold; a receiver’s drop rate; how often that defense catches interception-worthy throws; how often that offense and that defense convert on 3rd and 4th down. Judgment calls count too, at half, each weighted by how often the flagged player draws one. It doesn’t change the replay; it shows where the luck came from.

A worked example, from another game: Bears 24, Packers 22 Week 18 2024

The game
The Packers lost 24–22, but with both teams’ plays dealt out again, they won 99% of the replays. Packers: 100 × (0 − 0.99) = −99 Bears: 100 × (1 − 0.01) = +99
One play, scored on its own
C. Santos made a 51-yard field goal. Going by his record, the distance and the conditions, he makes that kick 62% of the time. From that distance a make is worth 4.5 points more than a miss, on average: the 3 points, plus the field position, since after a miss the other team takes over at the spot of the kick. The luck on the kick is how far the result landed from expected: what happened (1 for a make, 0 for a miss) minus the 62%, times those 4.5 points. That doesn’t make the kick lucky: a miss would have counted the other way, so a kicker who makes exactly his share nets zero over a season. The make: (1 − 0.62) × 4.5 = +1.7 points A miss would have been: (0 − 0.62) × 4.5 = −2.8 points Counted at the moment it came, with 0:02 left in the 4th quarter, that was worth 38% of a win to the Bears.
The season
The Packers finished 2024 at 11–6. The replays of those 17 games add up to 12.5 wins. Luck in wins: 11 − 12.5 = −1.5

Each step in full, and how well the numbers hold up: methodology.

The test

Luck, by definition, doesn’t repeat. Skill does. So there is a simple way to check a luck model without trusting it: stop a season after a few games, take the luck the model found out of every team’s record, and see whether that adjusted record forecasts what happens next better than the real one. If the model counted skill as luck, the adjusted record gets worse. If it counted luck as skill, or missed luck, it gains nothing.

The luck-adjusted record is the one in our adjusted standings: each game replayed with every kick, fumble bounce, 3rd and 4th down and judgment call redrawn at its expected rate, and the team credited with the share of replays it wins. Every forecast here is fitted on the other seasons and tested on the one left out, so no season grades its own homework.

We ran it on two records. The site’s full model covers 2019–2025 (1,871 games). That is too few seasons to separate forecasts this close (more below), so we also rebuilt the same luck from play-by-play for every season since 1999: 6,967 games and 861 team-seasons, with league-average expectations in place of each player’s own record.

Lucky teams fall back

After eight games, group the teams by how much luck is in their record. The actual record says the lucky teams are good. The adjusted record says they are closer to average. The rest of the season sides with the adjusted record.

Luck in wins after 8 gamesTeamsRecord so farLuck-adjustedRest of season
1.5 or more unlucky37.274.512.474
0.5 to 1.5 unlucky219.371.481.492
about even373.510.507.498
0.5 to 1.5 lucky192.621.506.507
1.5 or more lucky40.750.505.544

Win shares, 1999–2025. Luck in wins = wins minus luck-adjusted wins.

Teams that were 1.5 wins or more lucky had won .750 of their games while playing like a .505 team, and won .544 the rest of the way: 84% of the way back to their adjusted record. The unluckiest went from .274 to .474, 84% of the way. Every team drifts toward .500 after a hot or cold start, so the adjusted record doesn’t call the second half exactly either, but the teams whose records were luck drifted furthest, and in the direction their luck said.

The rest of the season

After n games, how well does each record predict the win share over the remaining games? The correlation is higher for the adjusted record at every one of these points, and a forecast built on it misses by fewer wins at every one of these points: by 1.49 wins on average, against 1.51 for the actual record.

AfterActual record, rLuck-adjusted, rMiss in wins, actualMiss, adjusted
3 games0.3270.3602.072.04
4 games0.3620.4031.911.88
6 games0.4040.4561.661.61
8 games0.4190.4581.411.38
10 games0.4470.4641.121.10
12 games0.3590.3900.920.90

861 team-seasons. Miss: mean absolute error of the forecast of a team’s remaining wins, fitted on the other seasons. Simply assuming a team keeps its pace misses by 2.73 wins after 4 games: most of the gain in any forecast is pulling a hot or cold start back toward .500, and taking the luck out adds to that.

The next game

The strictest test is every single game: forecast it from both teams’ records so far and score the forecast when it’s played.

Forecast fromPicks rightLog loss
Actual record so far61.8%0.645
Luck-adjusted record so far63.3%0.639

6,534 regular-season games from week 2, 1999–2025. Log loss scores the probability, not only the pick: lower is better, a coin flip scores 0.693. Difference, adjusted minus actual: −0.0067 ± 0.0021 (one standard error, resampling weeks).

Why 61.8% to 63.3% is a big step

1.5 points more games picked right looks small. It is a large gain, for three reasons.

  • It isn’t chance. Over 6,534 games the adjusted record called about 100 more right, and its log loss is lower by 3.2 standard errors. If taking the luck out changed nothing, chance alone would open a gap that large about once in 700 tries.
  • There is little room to gain. In the NFL a coin flip picks 50% of games, and even a forecast that knew both teams’ strength exactly would pick only about 68% (below). The actual record gets to 61.8%; taking the luck out wins 1.5 of the 6.2 points left: 24% of the remaining distance.
  • It uses nothing new. Other forecasts add last season, who starts at quarterback and the injury report. The adjusted record knows only the games the actual record already counts, adjusted for luck. The whole gain comes from reading the same games right.

How close this is to the best possible

The luck-adjusted record is one ingredient. Built into our full forecast, with who starts at quarterback, the week’s injury report and last season, it picks 65.3% of games (6,471 games, 2001–2025, each season forecast by a model fitted on the others; 65.0% since 2019, when the injury reports start). That is nearly as good as an NFL forecast gets, because so much of every game is luck. From how lucky the NFL is, measured on every game since 2019:

  • 22% of games are won by the team that played worse. Even someone who knew before kickoff exactly how both teams would play, snap by snap, would get those wrong, so a forecast that could see the future would top out near 78%.
  • Before kickoff the ceiling is about 68%. One game’s luck moves the final margin by about ±12.1 points; everything that separates the teams, by about ±7.7. A forecast that knew exactly how good both teams are, and only couldn’t know the luck, would pick about 68.0% of games. No forecast knows the teams that well, so every real one lands below it.
  • About 38% of the differences between records is luck. Over a season it moves a typical team’s record by about 1.9 wins. A forecast built on the record starts from something that is part noise, and that part is what the adjusted record takes out.
Picking winners, every NFL gamePicks right
A coin flip50.0%
Actual record so far61.8%
Luck-adjusted record so far63.3%
Our full forecast65.3%
Knowing both teams’ strength exactly (the ceiling before kickoff)about 68.0%
Knowing every snap in advanceabout 77.7%

Records: next games from week 2, 1999–2025. Full forecast: 2001–2025, from week 1. Ceiling: regular seasons since 2019, margin sd 14.3, one game’s luck sd 12.1. Bars start at 50%.

On that scale our forecast covers 85% of the distance from a coin flip to the ceiling, using everything known before kickoff and nothing from the betting market. The luck still to come in a game is what nobody can forecast, and it is what holds every forecast near two in three.

It gets stronger as the season goes on

Week by week, the forecast from the luck-adjusted record gets better, and its edge over the actual record grows.

WeeksGamesPicks right, actualAdjustedEdge over actual
Weeks 2–41,23656.7%58.8%+0.0058
Weeks 5–81,52060.7%62.4%+0.0045
Weeks 9–121,57363.8%64.9%+0.0096
Week 13 until the last two1,34864.7%65.7%+0.0056
Last two weeks85762.8%64.9%+0.0083

1999–2025, regular season from week 2. Edge over actual: log loss of the actual-record forecast minus that of the adjusted one (positive = adjusted better).

Against the actual record. The edge from taking the luck out is larger in the second half: +0.0079 in log loss from week 9 on, against +0.0051 in weeks 2–8. Two things grow as a season runs:

  • The luck piles up. After 3 games a typical team is 0.5 wins lucky or unlucky; after 12 games, 1.1. An actual record carries all of it into every forecast.
  • Forecasts lean harder on the record. In weeks 2–4 nobody’s record says much, so every forecast stays near a coin flip: they spread only ±9% around 50%. In weeks 9–12 they spread ±17%, and a lucky record moves them a lot. Taking the luck out is worth little while the record barely counts, and more once it does.

What the test changed

The same test graded our own settings, and three failed it. Counting every lucky break at +50% for momentum left the adjusted standings forecasting no better than the actual standings; without it they do better, so momentum is now off by default (the tests never found any). Judgment calls counted in full made a team’s luck start to repeat, which luck doesn’t: over 27 seasons they behave as about a third chance, and now count half. Injuries were valued a quarter too high against what they did to final scores, and now count three quarters. With those fixes, on the site’s own 2019–2025 games, the adjusted record picks 61.6% of next games right against 61.3% for the actual record.

What it doesn’t prove

  • That any single game’s luck number is exact. The test grades luck summed over many games, where errors cancel.
  • That seven seasons settle it. On the site’s full model alone (1,759 games) the difference in log loss is −0.0020 ± 0.0049, and the rest-of-season miss 1.58 against 1.60 wins: in the same direction, but inside the noise of that many games. The 27-season record is where the difference is measurable.

This week’s predictions · Every validation result · Adjusted standings 2026 · How lucky is the NFL?