v1 back to original
2026 World Cup · prediction markets · v2

The board, calibrated.

The same tournament model as the original board, with one change: its probabilities are temperature-calibrated to fix the tail-overconfidence an out-of-sample backtest exposed. The headline effect is on the extremes, where the raw model was too sure. Model vs market; the gaps are the calls, now honestly sized.
What v2 is, and why. Backtesting the model on the 2018 and 2022 World Cups (out-of-sample) showed it was overconfident at the extremes: confident calls landed about 60% of the time when the model said ~74%, and genuine upsets were priced near 0%. The standard fix is temperature scaling, one parameter T, fit on those past tournaments by minimizing log-loss, that pulls extreme probabilities toward the centre. Applied here, it is what collapsed an apparent favourite-longshot edge (e.g. backing longshots to advance) down to about ±2pp: that gap was the model's overconfidence, not a market mispricing. The picks are unchanged; what changes is how far the model is willing to stray from the price. It is a pre-committed fork run forward in parallel with v1, and it never edits v1.
Title race
Calibrated champion probability vs the market, top of the field.
● model (calibrated)    ● market  ·  the gap is the disagreement
Every forecast
Calibrated model vs market across all markets. Sort any column, filter by market, search a team. Edge = calibrated model − market.
loading…
,

Temperature. The single calibration parameter, fit out-of-sample on 2018 + 2022. T > 1 means the base model was overconfident, so v2 cools every probability toward the centre. Picks are unchanged; the edges shrink to their honest size.

What this fixed, and what it doesn't

The raw board showed a cluster of large edges on longshots to advance. The backtest says that was tail-overconfidence, not a priced edge: calibration pulls those toward the modest favourite-longshot bias the literature documents.

A caveat the project owns: the temperature is fit on single-match outcomes, which is exactly where it belongs (it drives the draw-calibrated group page). Applied to multi-match markets like reach-QF or champion it over-corrects, so treat the deep-run gaps as the model's added humility about who specifically goes deep, not as literal tradeable edges. The advance market, the closest to match scale, is the trustworthy one.

CLV and calibration grading stay on v1's pre-committed ledger (the auditable track record). This board is a calibrated current snapshot, not a separate graded ledger.