Market Research
Back to research
Predictions2026-07-01· 10 min

Prediction Market Calibration: Are Markets Efficient?

Testing Prediction Market Efficiency

We analyzed 2,400 resolved Polymarket contracts across politics, crypto, economics, and sports to answer a simple question: when Polymarket says something has a 70% chance of happening, does it actually happen 70% of the time?

Methodology

We collected resolution data for all Polymarket contracts that resolved between January 2024 and June 2026. We grouped contracts by their final trading price (our proxy for implied probability) and measured actual resolution rates.

Bins: 0-10%, 10-20%, 20-30%, 30-40%, 40-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-100%

A perfectly calibrated market would show a 45-degree line when plotting predicted probability vs. actual frequency.

Results

Predicted ProbabilityActual FrequencyN (contracts)Calibration Error

|----------------------|-----------------|---------------|------------------|

0-10%4.2%312-1.8% (well calibrated)
10-20%13.1%198-1.9% (well calibrated)
20-30%28.4%167+3.4%
30-40%42.7%224+7.7% ⚠️
40-50%53.1%356+8.1% ⚠️
50-60%61.8%298+6.8% ⚠️
60-70%68.2%245+3.2%
70-80%76.4%231+1.4% (well calibrated)
80-90%87.3%189+2.3% (well calibrated)
90-100%95.1%180+0.1% (well calibrated)

Key Finding: The Middle Is Overconfident

Markets are excellently calibrated at the extremes (<20% and >80%) but show systematic overconfidence in the 30-60% range. Events priced at 35% actually happen about 43% of the time.

Why? Three hypotheses:

  • Favorite-longshot bias: Bettors overweight unlikely outcomes (well-documented in sports betting), which pushes implied probabilities toward 50%.
  • Liquidity premium: In the 30-60% range, markets are most uncertain, leading to wider spreads and less efficient pricing.
  • Narrative bias: Events near 50/50 attract the most attention and debate, leading to price moves driven by narrative rather than information.
  • Category-Level Analysis

    Politics (n=680)

  • Most well-calibrated category overall
  • Slight overconfidence in 40-60% range (+5.2% average error)
  • Election markets were the best calibrated sub-category
  • Crypto (n=520)

  • Least well-calibrated category
  • Systematic overconfidence across all probability ranges
  • "Will X token reach Y price" contracts were the worst performers — events priced at 60% happened only 48% of the time
  • Economics (n=445)

  • Moderately calibrated
  • Fed rate decision markets were extremely well calibrated (within 2%)
  • GDP/inflation markets showed the overconfidence pattern
  • Practical Implications

  • Contrarian edge exists in the middle: If you can identify 35-45% contracts with genuinely higher true probabilities, there's a systematic edge.
  • Trust the extremes: Contracts priced at <15% or >85% are highly reliable signals.
  • Beware crypto prediction markets: These show the most systematic bias and are the least reliable for probability estimation.
  • Election markets are good: Among the most well-calibrated sources of political probability estimation available.
  • Limitations

  • Survivorship bias: we only analyze resolved contracts; many contracts are created and never gain liquidity
  • Sample period includes only 2 years of data
  • Polymarket-specific dynamics may not generalize to other platforms
  • Conclusion

    Prediction markets are remarkably efficient at the extremes but leave systematic edge in the middle probability range. For decision-making purposes, treat any Polymarket price between 30-60% with skepticism — the true probability is likely 5-8 percentage points higher than the market implies.