Methodology · college football
How college football luck is estimated, and how well it holds up
The same question as for the NFL: with the way both teams actually played, how often does each win a replay of the game in which every chance event happens at its expected rate? The difference between that and the result is luck. This page sets out what college football changes, the data work it takes, and every validation result.
1. What’s different in college football
The estimand, the luck index (100 × (result − replay win probability)) and luck in wins are the NFL’s, set out in full on the NFL methodology page. What changes:
- Who’s covered. Every game with an FBS team, from the 2025 season on, including games against FCS opponents. The engine reads FCS games too, for those teams’ records.
- The data. Play-by-play with expected points from CollegeFootballData.com via cfbfastR, merged from several stats providers. There is no charting of drops or interception-worthy throws and no injury report, so those kinds of luck are not counted. Rest was measured and has no effect in college football (−0.17 ± 0.11 points of margin per day of rest advantage between FBS teams, 4084 games), so it isn’t a component either.
- Our own win probability. The source’s win probability is computed in its own order of plays, which from 2025 is scrambled in many games (section 11), so we fit our own (section 5). It starts from the betting line: in college football a 40-point favorite’s early misses are not the swings they would be in a close game.
- Overtime. College overtime has no clock and its own rules. Win probability in overtime comes from a model of the format, per possession (section 5).
- Pooling within the game. The replay fills sparse situations from the same game rather than from a team’s earlier games (section 2): in college football an FCS team’s earlier games came against FCS opponents and would flatter it.
2. The replay with the same day’s form
Resampling. Each team’s offensive snaps in the game (runs, passes, and accepted penalties on them) are partitioned into nine strata by down (1st, 2nd, 3rd/4th) and distance (≤3, 4–7, 8+ yards). A replay builds drives by drawing each snap from the stratum of the current down and distance.
Pooling. With n same-game snaps in a stratum, a draw comes from them with probability n / (n + 12); otherwise from the team’s snaps in the same game at the same distance on any down, or all its snaps of the game. Garbage time is left out when at least 40 snaps remain (as college analysts define it: a margin over 43 points in the 1st quarter, 37 in the 2nd, 27 in the 3rd, 21 in the 4th), since a blowout’s backups don’t stand for how the game was played. We compared this with the NFL’s pooling (the team’s earlier games that season) and a version that uses only earlier games against the same division, on the 2024–2025 seasons with the calibration fitted on 2022–2023:
| Pooling | Holdout log loss | FBS vs FCS: replay → won | Predicts the rest of a season, r |
|---|---|---|---|
| Earlier games this season (the NFL’s) | 0.404 | 84% → 96% | 0.44 |
| Earlier games against the same division | 0.391 | 97% → 96% | 0.44 |
| This game only | 0.388 | 90% → 96% | 0.47 |
| This game, garbage time left out ✓ | 0.377 | 90% → 96% | 0.47 |
Chance at its expected rate. Interceptions at the league rate per throw (2.6%); fumbles lost at the league rate for the kind of play; field goals made at the college make curve for their distance; extra points at the college rate. College football has its own game: kickoffs start at the 25, punts net 38 yards, and teams go for it far more on fourth down than in the NFL (in the opponent’s half, 81% of 4th-and-1s and half of 4th-and-3s). The replay’s fourth-down policy follows what college teams do. A game tied after regulation counts half to each side: overtime is close to a coin flip.
Calibration. The raw replay is under-dispersed, and in college football it falls short of home field: home teams win more often than their plays alone replay to, most likely because part of home field works through what the replay draws at neutral rates (turnovers, kicks), and college home field is bigger than the NFL’s. A logit scale and a home shift correct both, π̂ = σ(T · logit π̂raw + h), with h only at the home team’s own stadium: T = 1.30, h = 0.28 (in an even game, 6.9 percentage points more for the home team), fitted by maximum likelihood on 2019, 2021, 2022, 2023 and evaluated on 2024–2025. 4,000 replays per game on the site.
3. Three reference estimates
- Before kickoff
- The closing line as a probability: Φ(s / 15.4), the 15.4 points being how far college final margins scatter around the line (games with an FBS team, 2019–2025). The NFL’s is 13.5.
- Same teams, fresh game
- The line plus this game’s luck-neutral efficiency, shrunk by its reliability (split-half r = 0.19): latent strength on the day has a standard deviation of about 8.0 points, against 13.4 for one game’s randomness.
- Only coin-flips re-rolled
- Every kick, fumble recovery, conversion and two-point try redrawn from its own probability, every judgment call given a fair coin for its direction, all else fixed.
4. Event-level luck
Each chance event is scored as outcome minus expectation, in expected points, ℓe = (oe − pe) · ve, and each play counts for one component at most, as in the NFL. The expectations and their shares:
| Component | Expectation | Share counted as luck | Game-to-game r | Predicts later margin, r |
|---|---|---|---|---|
| Kicks | Field goals against a logistic make model (distance as a linear spline with knots at 45 and 55 yards), shifted by the kicker’s own record; extra points against the college make rate (97.8%). | 100% | 0.05 | 0.02 |
| Fumble bounces | Which team recovers a loose ball, against the rate for that kind of play: the fumbling team keeps 50% on runs, 44% on passes and sacks, 50% on punts. | 100% | 0.02 | −0.06 |
| Judgment calls | Fouls that require an official’s judgment, valued at the expected points they transferred, including any play they nullified. | 0% | 0.19 | −0.10 |
| 3rd & 4th downs | 3rd- and 4th-down attempts against a logistic model of distance, down, goal-to-go and the offense’s early-down efficiency that day, shifted by both teams’ records; two-point tries against the college rate (46%). | 50% | 0.02 | 0.24 |
| Pre-snap flags | False starts, offside and similar fouls: discipline rather than chance where they repeat. | 100% | 0.15 | 0.02 |
| Special teams | Return touchdowns and blocked punts. | 0% | 0.13 | 0.17 |
Game-to-game r: correlation between a team’s average in its odd and even games of a regular season. Predicts later margin: its first-half average against its second-half point margin. For pure chance both should be near zero.
Kicker- and team-specific expectations, by empirical-Bayes shrinkage on strictly earlier plays (this season and the two before): kickers k = 80 (and 1000000000 beyond 45 yards), 3rd and 4th downs the offense k = 320 and the defense k = 320. Kickers are identified by team and surname, as the play text gives them.
5. Win probability and leverage
In regulation, the final margin from the offense’s point of view is normal around the score, the expected points of the current possession and what the betting line says the rest of the game is worth:
P(win) = Φ((d + 0.91 · EP + 1.23 · s · t/60) / (18.3 · √(max(t, 2.7)/60)))
with d the lead, s the line and t the minutes left (home field of 4.1 points where there is no line), fitted by maximum likelihood on 428,117 snaps of regulation. Against the source’s own win probability, in games where the source’s play order is sound (log loss on the result, lower is better):
| Season | Snaps | Ours | cfbfastR’s |
|---|---|---|---|
| 2024 | 110,433 | 0.352 | 0.433 |
| 2025 | 62,319 | 0.303 | 0.410 |
In overtime each team gets a possession from the 25; from the second overtime a touchdown must be followed by a two-point try, from the third each team only tries two-point plays. Win probability is computed per possession from the college outcome rates of overtime possessions (52% end in a touchdown, 21% in a field goal; 475 possessions) and the try rates, flat while a team has the ball and moving when the possession ends.
Leverage is as in the NFL: binary events take their realized swing (the two-point try and the extra point at the leverage after the score, where the source folds them into the touchdown), others their points at the local exchange rate φ(Φ−1(WP)) / (18.3 · √(t/60)), capped by the game state and by the play’s own swing.
6. Officiating
The source doesn’t say who was flagged. The play text does, in the provider’s words: the school (“Tulane Penalty, Holding”) or an abbreviation (“PENALTY TUL Holding”, “PENALTY on TULN-J.Smith”). Abbreviations are matched to teams through the yard lines they appear in across the season: on a run followed by the same offense’s next snap, “to the LAT 47” means 47 yards from Louisiana Tech’s goal line if LAT is Louisiana Tech. Where the text can’t tell and the play was wiped out, the flag is put on the team the ball moved against. The check: false starts are always the offense’s, and 98.4% or more land there in every season. A flag’s value and the treatment of judgment and pre-snap calls are the NFL’s; there is no per-player flag history in college data, so every flag counts in full.
7. Luck shares
The share of each component that behaves as chance is fitted as in the NFL: the one that makes a team’s luck-adjusted margin over the first half of its regular season best predict its raw margin over the second half, on 2019, 2021, 2022, 2023, from 1 and moving only for a gain above sampling noise. On 2024–2025 the luck-adjusted margin predicts later margins at r = 0.56, against 0.56 for the raw margin and 0.54 counting everything modeled as luck. As in the NFL, the luck-adjusted margin forecasts the second half no better than the raw margin: the shares count as luck only what can be taken out without making that forecast worse. 2020, played in a pandemic with shortened, conference-only schedules and opt-outs, is left out of fitting and validation.
8. Momentum
Across 38,008 pure-chance events worth a point or more (fumble bounces, kicks, judgment calls), the team that got the break did no better on its next possession than average: −0.013 ± 0.014 expected points per point of luck. As on the NFL pages, momentum is a setting: by default none, what the test measured; the panels can count every in-game luck event at up to +100% in the coin-flip replay and in luck in points and wins.
9. What the “adjust the assumptions” panel changes
It reruns the coin-flip replay and luck in points under your shares, momentum and choice of what counts. It doesn’t move the replay with the same day’s form: that replay draws kicks, bounces and turnovers at luck-neutral rates already.
10. Validation
5,312 games simulated (500 replays each); fitted on 2019, 2021, 2022, 2023, evaluated on 2024 and 2025.
When the replay says 90%, it happens about 90% of the time
Every game with an FBS team, 2019–2025 (2020 left out), grouped by the replay’s win probability for the home team: what it predicted against how often the home team actually won.
Show as a table
| Bin | Games | Predicted | Actually won |
|---|---|---|---|
| 0%–10% | 577 | 3.4% | 2.1% |
| 10%–25% | 415 | 17.6% | 22.7% |
| 25%–50% | 800 | 38.0% | 39.9% |
| 50%–75% | 966 | 63.2% | 59.7% |
| 75%–90% | 786 | 82.9% | 81.0% |
| 90%–100% | 1768 | 97.5% | 98.1% |
Holdout log loss on the same game’s result
Lower is better, but as in the NFL a low number is no virtue by itself: a replay built from a game’s own plays contains much of its result.
| Closing line | 0.478 |
| Replay with same day’s form, raw / calibrated | 0.386 / 0.377 |
| Coin-flip replay | 0.244 |
FBS against FCS
In 247 holdout games between an FBS and an FCS team, the replay gives the FBS team 92% on average; it won 96%. Mismatches come out a little closer than they are, so an FCS team’s near-miss shows as less of a heist than it was.
Does it predict the rest of a season?
Correlation between an FBS team’s value over the first half of its regular season and its win share over the second half (794 team-seasons).
| Replay wins (same day’s form) | 0.47 |
| Coin-flip wins | 0.47 |
| Closing line | 0.44 |
| Point differential | 0.48 |
| Actual wins | 0.45 |
Unbiasedness
Realized against expected rate of the favorable outcome for each binary component, every event 2019–2025, 2020 left out (z = the difference in binomial standard errors). Where |z| is above 2 the gap is 0.3 percentage points (third and fourth downs) and 0.2 percentage points (extra points): with this many events a small gap is detectable, and it is small against the luck in a game:
| Component | Events | Realized | Expected | z |
|---|---|---|---|---|
| Conversions | 154,351 | 43.6% | 43.9% | −2.7 |
| Conversions (2-pt) | 2,184 | 45.6% | 45.4% | 0.1 |
| Fumble luck | 9,672 | 47.8% | 47.1% | 1.5 |
| Kicking (FG) | 16,015 | 75.4% | 75.0% | 1.1 |
| Kicking (PAT) | 33,371 | 97.8% | 97.6% | 2.4 |
Market efficiency
Season-to-date luck and covering the next spread: r = −0.026 over 7,526 games.
11. Data quality
Two facts can be checked against things the source can’t get wrong together. Scoring: touchdowns, tries and field goals are read from the play text, and their points should add up to the final score. Order: the score never goes down, so a source whose plays run backward in time is out of order. We sort by the game clock and the running score, rebuild drives and expected points added from the plays’ own states, and read touchdowns, tries and fumble recoveries from the text rather than the provider’s labels, which contradict the text in some games.
| Season | Games | Scoring adds up exactly | Within 2 points | Source plays out of order | False starts on the offense |
|---|---|---|---|---|---|
| 2019 | 847 | 94.5% | 95.9% | 0.2% (42 games) | 98.9% |
| 2021 | 849 | 85.9% | 91.5% | 2.0% (120 games) | 98.4% |
| 2022 | 854 | 88.3% | 91.8% | 5.0% (308 games) | 98.4% |
| 2023 | 910 | 94.4% | 95.9% | 1.0% (129 games) | 99.7% |
| 2024 | 919 | 91.7% | 93.8% | 0.6% (117 games) | 99.6% |
| 2025 | 933 | 90.6% | 91.7% | 9.1% (814 games) | 99.6% |
| 2026 | 331 | 96.1% | 97.3% | 16.5% (329 games) | 100.0% |
Where scoring doesn’t add up, a play is usually missing from the source; a game whose plays account for less than 90% of its points says so on its page.
12. Data
Play-by-play with expected points, schedules and closing lines from CollegeFootballData.com through cfbfastR and the SportsDataverse releases; team details and further lines from ESPN via SportsDataverse. From 2025 many games’ plays come out of order in the source, and the source’s own win probability and expected points added were computed in that order, which is why we recompute both.
13. Limitations
- No charting of drops or interception-worthy throws and no injury reports: interception luck, drops and missing players aren’t counted.
- Onside kicks recovered by the kicking team, and fumbles on kickoffs, can’t be told apart reliably in the play text and aren’t counted.
- Kickers are identified by team and surname; a transfer starts a new record.
- Between an FBS and an FCS team the replay runs a little closer than the results (section 10).
- The season under way uses the model fitted on 2019–2025 until it is refitted after the season.
- A game appears once its play-by-play is published, usually the day after; plays missing from the source are missing here.