2026 World Cup · how the model is doing
The scorecard.
No cherry-picking: the running grade on every pre-committed forecast. Favourite accuracy and match log-loss against the uniform baseline, the v1→v2→v3 head-to-head as the draw fixes earn (or don't earn) their keep, Closing-Line Value, and the biggest surprises, the residuals the model didn't see coming. The honest version is the point.
Model versions, head-to-head
Each version is a pre-committed fork run forward in parallel, graded only on games both versions forecast. Log-loss and Brier, lower is better. v1 is never edited; the forks fork forward.
Model vs market, on calibration
The honest headline. Brier score over the 72 group-stage matches, lower is better, against the no-information baseline. The pre-committed model beats the baseline comfortably, and the real-money market beats the model.
Biggest surprises · the residuals
Played games ranked by how much the result surprised the model, the probability it gave the actual outcome before kickoff. Low = the model was confident and wrong (or a true upset). This is the residual the project is named for; the pattern in the misses is the finding.