How accurate is Saturday Bias?
Check your bias applies to us too. This page is our own report card, drawn straight from the backtest — no cherry-picked window, no rounding in our favor.
2026 — called live
Logged before kickoff, frozen at kickoff, graded as finals land. The backtest windows below are reconstructed after the fact — this section is the record we can't massage.
8 predictions logged for upcoming games — results land here as games finish.
Core window: 2005–2025
Every regular-season game with a prior-year rating, results, and recruiting data on file — the widest test we can honestly run.
Across 13,905 games from 2005–2025, our number missed the final margin by 12.6 points on average and picked the winner 74% of the time.
The naive baseline — last season's rating, regressed, and nothing else — missed by 14.0 points and picked 69% of winners. We beat naive on both — that's the bar we have to clear to ship.
SP+ missed by 11.1 points and picked 79% of winners. We lose to SP+ on average miss. SP+ is scored with full-season hindsight (CFBD publishes it as a season-final number), so it's a stretch reference, not the bar.
| Model | Avg miss (pts) | Win-pick rate |
|---|---|---|
| Saturday Bias | 12.6 | 74% |
| Naive baseline | 14.0 | 69% |
| SP+ (full-season hindsight) | 11.1 | 79% |
Calibration
When we say a team has an X% chance to win, how often do they actually win?
| Bucket | Predicted win% | Actual win% | Games |
|---|---|---|---|
| 0-10% | 6% | 6% | 502 |
| 10-20% | 15% | 13% | 835 |
| 20-30% | 25% | 25% | 1116 |
| 30-40% | 35% | 35% | 1405 |
| 40-50% | 45% | 46% | 1450 |
| 50-60% | 55% | 54% | 1587 |
| 60-70% | 65% | 66% | 1742 |
| 70-80% | 75% | 76% | 1736 |
| 80-90% | 85% | 85% | 1793 |
| 90-100% | 95% | 96% | 1739 |
Full window: 2015–2025
Adds returning production and staff continuity to the inputs above — those only exist for recent seasons, so this window is narrower.
Across 7,778 games from 2015–2025, our number missed the final margin by 12.7 points on average and picked the winner 73% of the time.
The naive baseline — last season's rating, regressed, and nothing else — missed by 14.1 points and picked 68% of winners. We beat naive on both — that's the bar we have to clear to ship.
SP+ missed by 11.0 points and picked 79% of winners. We lose to SP+ on average miss. SP+ is scored with full-season hindsight (CFBD publishes it as a season-final number), so it's a stretch reference, not the bar.
| Model | Avg miss (pts) | Win-pick rate |
|---|---|---|
| Saturday Bias | 12.7 | 73% |
| Naive baseline | 14.1 | 68% |
| SP+ (full-season hindsight) | 11.0 | 79% |
Calibration
When we say a team has an X% chance to win, how often do they actually win?
| Bucket | Predicted win% | Actual win% | Games |
|---|---|---|---|
| 0-10% | 6% | 8% | 289 |
| 10-20% | 15% | 14% | 476 |
| 20-30% | 25% | 24% | 642 |
| 30-40% | 35% | 38% | 757 |
| 40-50% | 45% | 45% | 817 |
| 50-60% | 55% | 54% | 894 |
| 60-70% | 65% | 65% | 953 |
| 70-80% | 75% | 75% | 959 |
| 80-90% | 85% | 85% | 970 |
| 90-100% | 95% | 96% | 1021 |
Fan arguments we tested
Every candidate we pre-registered before running the backtest — including the ones that failed.
| Argument | Effect (pts at 1σ) | Verdict |
|---|---|---|
| Trench mass diff Fans swear big lines win games. They're half right — big-line teams do win more, but only because good programs recruit big linemen. Once we account for talent and last season, line size adds nothing to the prediction. Across 11 seasons, no measurable effect on margin. | 0.42 | No effect |
| Staff continuity diff A new coordinator feels like it should swing a game. Across 11 seasons, the change doesn't move the final margin beyond what the roster already tells us. Shown as context, not a factor in our number. | 0.76 | No effect |
| Pace amplification More possessions should mean more chances for the better team to pull away — a fast game amplifying the mismatch. Tested against our shipped rating across 7 seasons: no effect. The edge is already the edge, regardless of pace. | -0.07 | No effect |
| Identity exploit net Run-heavy team meets a defense that can't stop the run — feels like extra points waiting. Across 7 seasons the matchup adds nothing beyond what the two ratings already say; if anything the point estimate leans the other way. | -0.29 | No effect |
| Boom bust edge Explosive offense against a havoc defense should be boom-or-bust. Across 7 seasons, the boom and the bust cancel: no measurable effect on the margin beyond the ratings. | 0.06 | No effect |