Brier score
Mean squared error between the market’s YES probability and the actual 0/1 outcome. 0 is perfect, 0.25 is a coin flip, 1 is consistently wrong. Lower is better.
Calibration
When a prediction market prices something at 70%, does it resolve YES about 70% of the time? This page tracks the answer — venue by venue, category by category — using the probability each market showed 24 hours before it resolved, not the near-certain price at settlement. That lead-time horizon measures genuine foresight, not convergence.
Markets priced at 70% should resolve YES about 70% of the time. Each dot is a probability bucket — on the dotted diagonal means the prices were honest; above it, the market was underconfident; below, overconfident. Dot size is sample count (1,214 graded markets).
| Priced at | Resolved YES (│ = priced) |
|---|
This page measures
This page does NOT measure
How much is a price worth a day, a week, a month before resolution? Each row grades the probability the market showed at that lead time. Expect Brier to worsen as the horizon grows — the gap quantifies how much late information is priced in. The at-settlement row is a baseline, not a forecast: it grades the last observed price, which for most markets was observed once the outcome was already known, so its near-perfect score measures convergence rather than foresight. Read the lead-time rows for skill.
| Segment | Resolved | Brier (lower = better) | Accuracy |
|---|---|---|---|
| At settlement (not a forecast) | 1112 | 0.081 | 87% |
| 24 hours out | 1135 | 0.100 | 84% |
| 7 days out | 540 | 0.120 | 81% |
| 30 days out | 115 | 0.117 | 83% |
Graded on the YES probability 24 hours before resolution. Read Brier first: directional accuracy flatters near-certain contracts (a market at 99% earns an easy “correct call” that says little about calibration quality), and per-venue samples differ in size and composition.
We select the latest valid snapshot strictly before each evaluation horizon — 24 hours, seven days, and 30 days before the event deadline. The settlement-adjacent price is kept only as a convergence baseline, not a graded forecast.
When the market settles YES or NO, that binary result becomes ground truth. No estimates, no backfill.
The falsifiable test of the Trust Score: if it means anything, High-trust markets should realize a lower Brier loss than Low-trust ones. Forward-only — scored from markets captured and resolved after tracking began (Jul 26, 2026), never backfilled.
| Trust tier | Resolved | Brier | Abs error | Large-error rate |
|---|---|---|---|---|
| High | 9 | 0.000 | 0.002 | 0% |
| Medium | 296 | 0.000 | 0.002 | 0% |
| Watch | 308 | 0.005 | 0.008 | 1% |
| Low | 1,556 | 0.053 | 0.117 | 16% |
| Overall | 2,169 | 0.039 | 0.085 | 11% |
Some tiers hold only a handful of resolved markets, too few to order High against Low — no verdict is claimed until every tier has ~10. This is a real result either way — the harness reports it honestly rather than assuming the score works.
Mean squared error between the market’s YES probability and the actual 0/1 outcome. 0 is perfect, 0.25 is a coin flip, 1 is consistently wrong. Lower is better.
The share of resolved markets where the side the market favored (above 50%) was the side that won. A blunter measure than Brier, but easy to read. Coin-flip (≈50%) predictions are excluded — they carry no directional signal.
Voided or non-binary settlements are excluded. Samples are small until enough markets resolve, so treat early numbers as directional. See the methodology for how scores are derived. ProbCast provides prediction-market data and analytics for informational purposes only. Not financial, trading, betting, investment, legal, or tax advice.
| Resolved YES |
|---|
| Markets |
|---|
| 0–10% | 1% | 419 | |
| 10–20% | 15% | 84 | |
| 20–30% | 23% | 66 | |
| 30–40% | 43% | 74 | |
| 40–50% | 45% | 87 | |
| 50–60% | 57% | 135 | |
| 60–70% | 62% | 45 | |
| 70–80% | 67% | 51 | |
| 80–90% | 86% | 43 | |
| 90–100% | 98% | 210 |
| Venue | Resolved | Brier (lower = better) | Accuracy |
|---|
| All venues | 1135 | 0.100 | 84% |
| Manifold | 595 | 0.171 | 72% |
| Polymarket | 479 | 0.022 | 97% |
| Kalshi | 61 | 0.023 | 98% |
Markets aren’t equally good at everything — the same venues can be sharp on macro and loose on politics. Thin segments (<10 resolved) are shown muted.
| Segment | Resolved | Brier (lower = better) | Accuracy |
|---|---|---|---|
| Macro | 636 | 0.127 | 80% |
| Crypto | 227 | 0.027 | 97% |
| Geopolitics | 133 | 0.068 | 88% |
| Ai | 70 | 0.127 | 79% |
| Politics | 69 | 0.129 | 81% |
The full analysis of this decay curve — what it says about how much foresight market prices actually contain — lives here.
Brier score and directional accuracy, per venue and per category. Lower Brier = better-calibrated prices.
ProbCast provides prediction-market data and analytics for informational purposes only. Not financial, trading, betting, investment, legal, or tax advice. ProbCast is not a betting platform, exchange, or broker and does not provide buy or sell recommendations. Intended for users 18+; availability of prediction-market activity varies by jurisdiction. Venue data is aggregated from third parties and may be delayed or inaccurate. See our Terms of Use.
Market-implied probabilities can be wrong, illiquid, manipulated, or affected by ambiguous resolution rules.
© 2026 ProbCast · [email protected]