Working draft · 2026

Cross-Venue Price Discovery in World Cup Prediction Markets

Prabhat M · portfolio · repo · pre-registration
This note is complete: the 2026 World Cup is over and every result is final. The lead-lag event study and the Hasbrouck/Gonzalo-Granger decomposition agree that Polymarket leads price discovery; the harvestability forward-test shows the lead is un-harvestable at size; and the pre-registration was graded in public after the final, 6 pass, 2 fail, 3 inconclusive. Every empirical claim is bound to code and to a pre-registration committed before kickoff, so the results are falsifiable, not flexible.

Abstract

Two large real-money prediction markets, Kalshi (US, regulated) and Polymarket (global, on-chain), priced every 2026 World Cup outcome continuously, tracked against the global bookmaker consensus. Using millisecond-resolution order-book and trade captures across both venues, plus an on-chain trade-flow layer reconstructed from Polymarket's CTF Exchange, we study where price is discovered: which venue moves first when information arrives, and whether that leadership flips between the quiet pre-match regime and the high-information in-play regime. We decompose each cross-venue quote into belief and margin and find the widely-quoted "5 to 8 cent" inter-venue gap is almost entirely the house margin: de-vigged, the two prediction markets agree to ~0.15pp on the title race, and a relative-value convergence trade returns a documented loss net of costs. The residual belief gap is small but structured by audience (a home-crowd tilt). The central price-discovery result, from cross-correlating mid-changes around each goal shock, is a clean positive: across 392 goal shocks over 86 matches Polymarket leads the repricing a median +600ms ahead on a 72% share of events (Polymarket leads 281, Kalshi 111), and the formal decomposition agrees, a Gonzalo-Granger permanent-component share near 79% across 31 cointegrated matches, Polymarket leading 30 of 31. Two independent estimators name Polymarket the price leader, both computed on de-vigged mids so the result survives the ~59% order-direction problem. We then ask the question the lead invites, and answer it with two disclosed forward-tests: the cross-venue gap is not a harvestable edge. The pre-match convergence trade is a clean null (a cost illusion), and the in-play lead-lag is un-harvestable after the cost of immediacy: at the goal the book withdraws, best-price depth collapses to ~0.5% of normal, spread ~8× wider, so the stale-quote gap is 0% capturable once gated on the depth actually resting. That collapse-and-refill is the maker pulling quotes against toxic, information-motivated flow, adverse selection observed in real time, not slow pricing. The one corner where a mechanical view plausibly beats the human-driven price is the favorite-longshot bias at price extremes. The through-line: price discovery here is real and measurable, and almost none of it is harvestable, the honest pro-market reading.

Contribution. The cross-venue, pre-match-vs-in-play price-discovery comparison across the regulated and on-chain prediction-market venues is, to our knowledge, the named-but-unexplored question in the recent prediction-market microstructure literature; everything else here is a careful replication or a methods contribution (the confederation-shrinkage baseline correction, 4.3, and the on-chain trade-direction layer, 3.2).

1. Introduction

Real-money prediction markets aggregate dispersed information into a single live probability. The 2026 World Cup is a natural experiment: a 39-day, 104-match global event priced simultaneously by a US-regulated venue (Kalshi) and a global on-chain venue (Polymarket), against the global bookmaker consensus, with a dense, exogenous, precisely-timed information stream (goals). The recent literature names a gap: cross-venue price discovery has not been measured, in particular whether discovery leadership differs between the low-information pre-match window and the high-information in-play window. This note answers that question. We ask three things, in order of novelty:

Framing held throughout: the market is the subject, not the opponent. Where our independent model disagrees with a liquid price, the prior is that the model is wrong, and we test that prior against a third source before claiming an edge (4.3).

2. Related work

Hasbrouck (1995) information share and Gonzalo-Granger (1995) component share are the standard tools for attributing price discovery across venues trading one asset; we apply them to de-vigged mid-prices to sidestep the order-direction problem (3.2). Croxson and Reade (2014) show betting prices update efficiently and near-instantly to goals, with no systematic drift, the in-play benchmark our event study (5.3) builds on. Snowberg and Wolfers (2010) decompose the favorite-longshot bias into risk-love vs misperception; we measure its strength by probability decile and compare books vs prediction markets (5.4). Bürgi, Deng and Whelan (2025) document the same bias in 300,000+ Kalshi contracts, and the SoK of Rahman, Al-Chami and Clark (2025) maps modern (incl. on-chain) prediction-market microstructure; neither measures cross-venue price discovery, the gap this note fills.

3. Data and infrastructure

3.1 Capture

A 24/7 collector logs Kalshi and Polymarket order books and the bookmaker/exchange lines on a fixed cadence, and, per marquee match, a millisecond WebSocket capture records every book and trade message from both prediction venues to an append-only tape with a local-clock timestamp on each event. A dry-run capture recorded ~172k events at a 6ms median inter-event time, far finer than the tens-of-seconds scale on which a goal reprices, the regime where lead-lag is identified. Connection-control events (connect / disconnect / sequence gap) are logged in-band so the analyzer masks venue outages rather than mistaking them for quiet markets.

3.2 The order-direction problem and the on-chain layer

Inferring trade direction from a public WebSocket feed is unreliable (a documented ~59% classification ceiling), which corrupts any order-flow-imbalance signal. We avoid it two ways: (i) all price-discovery estimation uses mid-price moves, which are direction-agnostic; and (ii) for true signed flow we read Polymarket's CTF Exchange OrderFilled events directly from the chain (Polygon RPC / subgraph), giving exact taker direction.

4. Methods

4.1 De-vigging

Implied probabilities are recovered from quotes by multiplicative de-vigging (removing each venue's overround), so cross-venue comparisons are belief-to-belief, not quote-to-quote. The closing (last pre-kickoff) quote is the calibration forecast.

4.2 Price discovery

On the synchronized de-vigged mid-price series for a contract across venues, we estimate the Hasbrouck (1995) information share (with the standard bounds) and the Gonzalo-Granger component share, computed separately for the pre-match and in-play windows. The event study (5.3) classifies each goal by surprise (pre-goal win-probability of the scoring side) and measures the reaction path: overshoot magnitude and mean-reversion half-life.

4.3 The independent baseline and its correction

An Elo-plus-squad-value goal model serves as an independent reference, not a competitor to the market: it lets us ask "where do model and market disagree, and who is right" against a third source (de-vigged bookmaker match odds). A methods contribution falls out of this: raw Elo inflates near-disconnected confederations (they mostly play themselves), so we apply an empirical-Bayes confederation shrinkage estimated from inter-confederation "bridge" games and scaled per team by global connectivity. It validates out-of-sample (+4.6% cross-confederation ranked-probability score, Diebold-Mariano p ≈ 0.009; within-confederation untouched as a placebo).

5. Results

5.1 Law of one price holds; the visible gap is margin (final)

De-vigged, Polymarket and Kalshi title prices agree to ~0.15pp on average across the 48-team field; the largest standing gap is England (~1pp). The "5 to 8 cent" gap the press quotes is mostly the house margin: Kalshi's overround runs ~5.4% vs Polymarket's ~3.0% (~1.8x), so the durable venue difference is cost, not price. The small surviving belief gap is structured by audience: the American book is richer on USA, Mexico, Netherlands; the global book on England, Portugal, Japan, Brazil, a home-crowd tilt. A liquidity asymmetry underlies this: Polymarket quotes roughly 27x the depth of Kalshi at the same spread on the title market.

5.2 Cross-venue price discovery: Polymarket leads in-play (matured)

On the first marquee captures (Mexico-South Africa, Qatar-Switzerland, Brazil-Morocco, Haiti-Scotland), the in-play lead-lag is estimated by cross-correlating binned mid-changes in a window around each auto-detected goal shock. A quality gate keeps only events with genuine positive co-movement (best cross-correlation ≥ 0.5) and a plausible lag (≤ 8s), discarding spurious detections: a "16-second lead at r = -0.70" is two books moving oppositely, a stale-tick artifact, not price discovery. An early four-match read (n=8) put Polymarket a median +500ms ahead, leading 6 of 8 with an interquartile range entirely positive, in the pre-registered direction (P6: the deeper-liquidity venue leads; Polymarket quotes ~27x the depth), but under the pre-registration's decision rule (evaluate at n ≥ 15-20 cross-venue matches) a four-match flash stays a non-claim, so it was never posted. It held. As the pool grew to 392 clean goal shocks over 86 matches, Polymarket leads a median +600ms ahead on a 72% share of events (Polymarket leads 281, Kalshi 111; interquartile range [0, +800]ms). The formal decomposition agrees and sharpens it: estimating the Hasbrouck (1995) information share and the Gonzalo-Granger (1995) permanent-component share on the de-vigged mids, Polymarket's component share is near 79% across 31 cointegrated matches, leading in 30 of 31. Two independent estimators, a model-free event study and a cointegration decomposition, both name Polymarket the price leader, and because both run on mids the result sidesteps the ~59% order-direction problem (3.2). The goal-shock event study (5.3) and the formal pre-match-vs-in-play split fill in as more tapes land; each match is archived per game so the sample is auditable, not overwritten.

Who prices a goal first, in-play
Two independent estimators · both on de-vigged mids · tournament-final sample
Goal-shock event study — share of 392 decisive shocks
Polymarket72%
Kalshi28%
Gonzalo-Granger permanent-component share — 63 cointegrated matches
Polymarket81%
Kalshi19%
Polymarket leads the repricing a median +600ms ahead (281 shocks to 111) and holds the majority permanent component in 61 of 63 cointegrated matches. Both estimators run on mid-prices, so the result sidesteps the ~59% order-direction problem.

5.3 Goal-shock event study: the market under-reacts (suggestive)

Conditioning each goal on a calibrated clock-and-Poisson fair-value benchmark anchored to the pre-match price (Croxson-Reade sets the efficiency prior: goals reprice near-instantly, no drift), the market's post-goal move is a median 0.55× the model's fair jump in log-odds, and it under-shoots the fair move in 7 of 8 cleanly-reconstructed matches. The move sticks (60-second reversion ~0), so this is a persistent under-reaction, not an overshoot. A decisive check rules out the alternative that the benchmark merely over-reacts: the model's larger post-goal probability is the better forecast of the eventual result (Brier 0.073 vs 0.113; log-loss 0.235 vs 0.361), so the larger jump was warranted. Reported as suggestive, not firm: the cluster-honest unit is the match, and 7 of 8 is a sign-test p ≈ 0.07. The raw-probability "value of a goal" curve and an apparent "underdog goals move twice as far" both dissolve once the log-odds geometry is removed, and are not claimed.

The market under-reacts to the goal it just priced
Post-goal move vs a calibrated fair-value benchmark, in log-odds · 8 clean-reconstruction matches
Model fair jump1.00×
Market move0.55×
The market books just over half the fair repricing and stays there — a persistent under-reaction the model's larger jump was right to make (Brier 0.073 vs 0.113 on the eventual result). The lag is a real inefficiency; whether it is tradeable is §6.2's question.

5.4 Calibration and the favorite-longshot bias

The favorite-longshot bias is visible pre-tournament in the 1-cent tick structure of longshot contracts. The half-spread by probability decile and the books-vs-prediction-markets comparison are a descriptive replication (Snowberg-Wolfers, 2010). The graded verdict (CORP reliability, Brier decomposition, slope), pre-registered as P1 (the markets are well-calibrated), is now final over 72 group-stage matches: the market is the best-calibrated forecaster, a Brier skill of +23.6% over the no-information baseline with a near-perfect reliability slope of 1.07 (1.0 is ideal). The pre-committed model trails it, +21.0% skill on a slope of 0.87 — mildly over-confident at the extremes, the favorite-longshot signature — which is exactly the pro-market premise: a real-money price out-calibrates a transparent fundamental model.

Reliability: predicted probability vs what happened
72 group-stage matches · points sized by sample · closer to the dashed 45° line = better calibrated
pre-committed model · Brier 0.503 · slope 0.87market · Brier 0.487 · slope 1.07perfect calibration
The market curve tracks the diagonal; the model bows below it at the high end (backing favorites a touch too hard). Both are close, and the gap between them is not significant: paired on the same 72 group-stage matches, the market's 0.487–0.503 Brier edge carries a bootstrap 95% CI of [−0.011, +0.044] (paired-t p = 0.25) and it scores better in only 34 of 72. P1 passes on its pre-registered point rule; the honest reading is that the market is well-calibrated, not that it beats the model.

The favorite-longshot bias is also the one corner where a mechanical view plausibly beats the human-driven price, and the only market-facing position the project takes. The independent baseline's advance probabilities are systematically more extreme than the market in both directions (favorites priced higher, longshots lower), the signature of the bias that persists even in deep prediction markets at the contract-price extremes. A model carries no psychological longshot premium, so its extremeness points the exploitable way. This underwrites a small, diversified basket in the paper track record: fade the overpriced longshots, back the underpriced favorites, sized as the modest systematic tilt it is, not single-name conviction. Whether it is a real edge or model tail-overconfidence is itself a calibration question, graded after the group stage. Notably it is the advance market that carries the signal (near-zero margin), whereas the reach-round ladder is 12 to 31% overround, where a model's apparent "fades" are the vig, not an edge.

6. Two disclosed forward-tests, two nulls

The two ways the cross-venue gap might be a harvestable edge, each tested out-of-sample rather than asserted, each disclosed with its rule.

6.1 Pre-match convergence: a cost illusion (final)

The law-of-one-price result (5.1) predicts there is no convergence arbitrage to harvest, and we tested that out-of-sample rather than asserting it. Rule: when the de-vigged Polymarket-Kalshi belief gap on a title widens past 1.0pp, go long the cheap venue and short the rich one, exit on convergence below 0.3pp or after a horizon, net of a 0.5pp modeled round-trip cost. Buildup result: 6 trades, -2.6pp total, 0% hit rate, per-trade Sharpe -1.95. The gap is real but does not converge enough to clear costs: the visible "edge" is a cost illusion, exactly what law-of-one-price implies.

6.2 In-play lead-lag: un-harvestable after the cost of immediacy (firm)

The 5.2 lead means Kalshi reprices a goal slightly behind Polymarket, so the natural follow-up is whether that lag is capturable. I answer it with a cost-of-immediacy ledger over 405 goals in 66 matches: at the instant Polymarket reprices, the median gross gap to Kalshi's stale quote is 12.0¢, a follower pays a median 1.4¢ to take Kalshi's posted price, and the median net is +10.8¢ on paper. Each is a median across matches of that match's median goal — the match is the unit, and the three are computed independently, so they are not meant to subtract. That paper number is a trap, and surfacing it is the point. At the goal, Kalshi's best-price depth collapses to ~0.5% of normal, the spread blows out ~8× on Polymarket and ~2× on Kalshi, and the book takes ~3–4s to refill. There is no resting size to hit at the stale price: by the time depth returns, the quote has caught up. So of the +10.8¢ paper gap, the median match yields no harvestable goal — a median across the 66 ledger matches, which says the typical match offers a follower nothing, not that no goal was ever capturable (goal-weighted on the 21-match reconstructible subset: 11.1%, clustered in 6 of those 21, and unidentifiable in advance). This is a liquidity-withdrawal result, not slow pricing: the collapse-and-refill is the market-maker pulling quotes against toxic, information-motivated flow, adverse selection observed in real time. Consistent with the high-frequency lead-lag literature, where the versions that do profit (Poutré, Dionne and Yergeau, 2024) require colocation and limit-order execution a read-only, paper-only study cannot access. A real lead, un-harvestable net of the cost of immediacy, which is to say not an edge for anyone trading across the two books.

A real lead, zero harvestable edge
Cost-of-immediacy ledger at the goal · 405 goal shocks across 66 matches · medians across matches
Gross stale-quote gap+12.0¢
Cost to lift Kalshi−1.4¢
Net on paper+10.8¢
Actually harvestable0%
The +10.8¢ paper edge is the trap, and surfacing it is the point. At the goal Kalshi's best-price depth collapses to ~0.5% of normal and the spread blows out ~8×, refilling in 3–4s — there is no resting size to hit at the stale price, so once gated on the depth actually resting, 0% is capturable. That is a maker pulling quotes against toxic, information-motivated flow — adverse selection observed in real time, not slow pricing.

7. Pre-registration and grading

Eleven falsifiable, dated predictions were committed to a tagged git commit before kickoff, with binding methods, named primaries, and PASS/FAIL/INCONCLUSIVE decision states under proper scoring rules. The two primaries are P6 (cross-venue lead-lag: the deeper venue leads) and P1 (the markets are well-calibrated); both pass. Graded publicly on 2026-07-19 by scripts/grade_prereg.py: 6 pass, 2 fail, 3 inconclusive. Both failures are results in their own right — P3 is the raw law-of-one-price gap that de-vigging explains as margin, P10 the goal-overreaction edge already arbed away — and the three inconclusives are data-forced and recorded in the pre-registration addendum.

8. Discussion and limitations

The unifying finding is a discipline for telling real edges from mirages. The cross-venue gap was probed three ways. The pre-match convergence trade is a cost illusion (6.1, a clean null). The in-play lead-lag is a real lead (5.2) that is un-harvestable after the cost of immediacy (6.2: at the goal the book withdraws to ~0.5% depth, so 0% of the stale-quote gap is capturable). The favorite-longshot wedge (5.4) is a real but modest systematic tilt, the lone position the project takes. Two of the three look like alpha in a frictionless backtest and are not; the methods that separate them, de-vigging before calling any gap, crossing the real bid/ask, and reading the depth at the event, are the contribution as much as any single number. Price discovery here is genuine and measurable, and almost none of it is harvestable, the honest reading and the pro-market one.

The headline is pro-market. Across a five-layer model-vs-market scan of 238 contracts, the liquid winner market is efficient and the only soft corners are thin, new, or structural:

layerdepthmean abs gapverdict
Winnerdeepest / most liquid0.4ppefficient; model agrees with the sharp market
Advance (R32)liquid4.3ppour model under-rating minnows, not market softness
Reach QFthinner3.3ppthin-market softness (partly model tilt)
Reach SFthin1.9ppthin-market softness (partly model tilt)
Champion (elim.)thinnest / newest1.0ppfavorite overpricing, partly model tilt

Where our model disagreed most (minnow advancement, isolated confederations) an independent bookmaker sided with the market and the error was ours to fix (4.3). Limitations, stated plainly: single-tournament sample; the calibration n (72 group-stage matches — the pre-committed ledger does not reach the knockout rounds) supports a wide slope band and cannot separate market from model (p = 0.25), so the calibration claims stay humble; the study is read-only and paper-only (no orders, no capital); and the in-play results depend on clean marquee-match captures, the data-quality risk we actively manage.


References