ProbCast Research · living analysis, updated 2026-09-06
How far in advance are prediction markets right?
“Prediction markets are accurate” is usually measured the easy way: look at the price the moment before a market resolves. By then the outcome is mostly known and the price has converged to near 0 or 1 — a near-perfect score that measures certainty, not foresight. The interesting question is how good the price was before the outcome was obvious.
ProbCast captures every tracked market's probability at fixed lead times before it resolves — then grades those frozen predictions against what actually happened. Thispage recomputes every time markets settle, from the 1,303 resolved outcomes captured across Polymarket, Kalshi, and Manifold.
Each row below grades only the outcomes for which a price was actually recorded at that lead time, so the counts are far smaller than the archive: a market we first saw at settlement carries no forecast to grade, and a price taken after resolution is not a prediction. Read the graded-forecast counton each row as that row's real sample size.
The decay curve
| Prediction taken | Graded forecasts | Brier (lower = better) | Directional accuracy |
|---|---|---|---|
| At settlement (not a forecast) | 1,115 | 0.080 | 87% |
| 24 hours out | 1,140 | 0.100 | 84% |
| 7 days out | 542 | 0.120 | 81% |
| 30 days out | 116 | 0.118 | 83% |
Reading the table: a Brier score is the mean squared error between the predicted probability and the outcome (0.25 is what always guessing 50% scores; 0 is perfect). Directional accuracy is how often the >50% side won. Coin-flip prices within ±2pts of 50% are excluded from grading — they make no call to grade.
What this means
The pattern is consistent: prices at resolution look near-oracular, a day out they are good but clearly imperfect, and the further back you go the closer they drift toward an informed guess. That decay is the honest answer to “do prediction markets work?” — they aggregate information impressively fast near the event, and the premium for reading them early is real but bounded.
It is also why ProbCast grades venues at a fixed 24-hour lead time on the accuracy leaderboard instead of at resolution: any venue can look perfect at settlement; foresight is what differentiates them.
Caveats, honestly stated
- Samples shrink as lead time grows: a market must have existed (and been tracked) N days before resolving to be graded at horizon N, so long-horizon rows are smaller and noisier.
- The sample skews toward short-dated markets (Kalshi's daily and weekly series resolve constantly), so composition differs across horizons — this is a living measurement, not a controlled experiment.
- The sample currently clears ProbCast's credibility gate (size, prediction spread, both outcomes observed). Numbers on this page move as new outcomes settle; that is the point.
ProbCast provides prediction-market data and analytics for informational purposes only. Not financial, trading, betting, investment, legal, or tax advice.