
We Checked “The Market Already Knows” Against 2.2 Million Resolved Contracts
The Brier Score Math Behind “Prediction Markets Are Smarter Than Experts”
DN Market Structure Series
"Prediction Markets Are Smarter Than Experts," Checked Against the Actual Brier Scores
Polymarket and Kalshi are routinely described as oracles that outperform pundits, pollsters, and professional forecasters. Two different, competing academic studies of the same 2024 election reached opposite conclusions about which platform actually won. Includes the DN Prediction Market Calibration Scanner.
DECENTRALISED NEWS · MARKET STRUCTURE & FORECASTING SERIES · SEPTEMBER 2026
"The market already knows" has become one of the most reflexively repeated lines in crypto and political commentary, usually deployed to shut down an argument rather than open one. Polymarket and Kalshi have earned a reputation, largely through 2024 election coverage, as something close to an oracle: aggregate the wisdom of thousands of traders with real money on the line, and you get a probability more honest than any pundit, pollster, or professional forecaster could produce. The claim has a real, quantifiable, testable version, forecasting researchers have been measuring exactly this for over a decade using a metric called the Brier score. It also has a much shakier popular version, and the two get conflated constantly, including by the platforms themselves.
This piece tests the claim properly, using the actual published calibration research on both platforms, benchmarked against the gold-standard academic measure of human forecasting skill: the IARPA-funded Good Judgment Project's superforecasters. The honest answer is more interesting than either "markets are magic" or "markets are hype."
DN AI Summary
Two genuinely different metrics get conflated whenever someone claims prediction markets are "smarter than experts": calibration (does a market saying 70% actually resolve Yes about 70% of the time, measured by the Brier score) and hit rate (what share of individual calls turned out correct). On calibration, both major platforms perform well: a 2.24-million-market Kalshi study found Brier scores falling from roughly 0.08-0.09 at a three-month horizon to about 0.02 at close, and Polymarket's own cited Dune analysis reports a Brier score around 0.084 across resolved markets, both comfortably better than a random-guessing baseline of 0.25. But a December 2025 Vanderbilt University study of the actual 2024 election found Polymarket, the platform most credited with "calling" that race, had the worst hit rate of the three major venues at 67%, behind Kalshi (78%) and PredictIt (93%), the lowest-volume, most tightly position-capped platform of the three. Benchmarked against the Good Judgment Project's elite superforecasters, whose published Brier scores (often below 0.12 on the tournament's original doubled scale, roughly 0.05-0.07 converted to the standard 0-to-1 scale most market studies use) were also over 30% more accurate than US intelligence analysts with classified access, prediction markets are competitive with the best human forecasters in aggregate, but the highest-volume, most-cited platform is not obviously beating the most disciplined human forecasting process, and was outperformed by a smaller, stricter rival on the single highest-profile test case available.
Two metrics, constantly conflated
The Brier score, developed by meteorologist Glenn Brier in 1950 and adopted as the standard forecasting-accuracy metric ever since, is the mean squared difference between a predicted probability and the actual outcome (1 if it happened, 0 if it didn't). A perfect forecaster scores 0. Someone who always guesses 50/50 on a binary event scores 0.25. A forecaster who is confidently and consistently wrong can score close to 1, worse than a coin flip. This is the metric nearly every rigorous prediction-market study cites, and it measures something specific: whether a stated probability, averaged across many predictions, matches reality.
Hit rate measures something different and much simpler: out of a set of binary calls, what share turned out correct. A market or forecaster can have excellent calibration and a mediocre hit rate on any single high-profile subset of questions, particularly a small number of close, contested, thinly-traded races, exactly the kind that get the most media attention. This distinction is not a technicality. It is the entire explanation for why two credible academic studies looked at the same 2024 US election and reached apparently contradictory conclusions about which prediction market "won."
A working paper by Nicole Kagan and Rubens Baiocchi examined 2,243,741 resolved Kalshi markets across 11 categories from the platform's 2021 launch through mid-2026, the largest study of its kind. It found Kalshi prices were "extremely well calibrated" in aggregate as markets approached resolution, with Brier scores falling from roughly 0.08 to 0.09 at a three-month horizon to about 0.02 at close, and naive accuracy rising from 88.3% three months out to 97.2% at close. Kalshi's own research team, in a separate August 2026 release covering more than 2.2 million data points, reached a similar conclusion: the platform's forecasts are closely correlated with how often events actually occur. Polymarket's own accuracy page cites a public Dune dashboard, built by data scientist Alex McCullough, reporting prices over 90% accurate a month before resolution, 95 to 96% accurate in the final hours, and a Brier score around 0.0838 to 0.0843 across resolved markets.
Both figures comfortably beat the 0.25 random-guessing baseline, and both are genuinely impressive at the scale involved, millions of resolved contracts across categories ranging from elections to weather to corporate earnings. If "prediction markets are well calibrated in aggregate" were the entire claim, it would hold up cleanly against the data.
The 2024 US election is the single event most responsible for prediction markets' "smarter than experts" reputation, and it is also the event most directly studied for hit-rate accuracy. A December 2025 Vanderbilt University study, and a separately cited 2025 academic analysis covering $2.5 billion in political-market volume, both found the same ranking: PredictIt correctly called roughly 93% of political markets, Kalshi 78%, and Polymarket, the platform that received by far the most media coverage for "predicting" the election, just 67%, worse than both of its smaller rivals.
| Platform | 2024 election hit rate | Position limits | Trading volume |
|---|---|---|---|
| PredictIt | ~93% | Strict ($850 cap) | Lowest of the three |
| Kalshi | ~78% | CFTC-regulated, higher caps | Mid-to-high |
| Polymarket | ~67% | Uncapped, crypto-native | Highest, most media coverage |
The proposed explanation in the research is structural, not random noise. PredictIt's strict per-trader position limit prevents any single well-capitalized actor from dominating the order book, forcing something closer to genuine crowd aggregation. Polymarket's uncapped, whale-friendly structure allows large traders to move prices on thin, low-liquidity contracts, sometimes for reasons unrelated to superior information, which can distort the implied probability on exactly the highest-profile, most-watched markets, even while the platform's overall, volume-weighted calibration across millions of smaller contracts still looks strong.
The benchmark almost nobody checks: how good are actual expert forecasters?
The IARPA-funded Good Judgment Project ran a four-year forecasting tournament from 2011 to 2015, pitting roughly 25,000 volunteer forecasters, including domain-expert teams, against each other on hundreds of geopolitical questions. The top 2% of participants, the "superforecasters," consistently posted Brier scores below 0.12 on the tournament's original scoring convention and were more than 30% more accurate than US intelligence analysts working the same questions with access to classified information. That original IARPA convention uses a doubled scale where guessing 50/50 scores 0.5, rather than the 0.25 baseline used in most modern prediction-market studies; converted to the same standard scale Kalshi and Polymarket report on, published superforecaster scores in the 0.10 to 0.14 range translate to roughly 0.05 to 0.07.
Measured on the same footing, that puts elite human superforecasters in a similar or slightly better range than Polymarket's aggregate 0.084 and comparable to Kalshi's mid-horizon figures, though Kalshi's near-resolution Brier score of roughly 0.02 pulls ahead of the converted superforecaster range. The honest reading is not "markets beat experts" or "experts beat markets." It's that the very best human forecasting processes and the largest prediction markets are operating in a similar performance tier overall, and markets' real, measurable edge shows up specifically in the final stretch before resolution, when new public information needs to be priced in fast, a genuine structural advantage over a forecasting tournament that updates on a slower cadence, not evidence that a market's day-one or month-out price is inherently wiser than a trained human's.
- Kalshi's 2.24-million-market study and Polymarket's own cited Dune dashboard both show genuinely strong calibration, comfortably beating the random-guessing baseline at scale
- Markets clearly excel in the final hours before resolution (95%+ accuracy), a real structural edge over slower human forecasting processes at impounding late-breaking information
- Even Polymarket's "worst" 2024 hit rate of 67% is well above a coin flip, and the platform's own calibration data across a far larger sample than the single 2024 election looks strong
- Prediction markets operate continuously across thousands of live questions simultaneously, a breadth no team of human superforecasters can match
- On the single highest-profile real-world test available, the 2024 election, the most-cited platform (Polymarket) had the worst hit rate of the three major venues, behind both Kalshi and the far smaller PredictIt
- Structural position-limit differences suggest whale-driven price distortion on exactly the thin, high-attention markets people cite prediction markets for, not the deep, liquid ones the aggregate calibration stats are drawn from
- Converted to a consistent scale, elite human superforecasters are competitive with or better than Polymarket's aggregate calibration, undercutting the specific claim that markets are smarter than the best available experts
- Calibration and hit rate are routinely conflated in popular coverage, and the platform average, aggregate calibration figure is not the number that was actually tested against experts on the marquee 2024 case
The tool: computing calibration for yourself
The DN Prediction Market Calibration Scanner below computes a real Brier score from any set of predicted probabilities and outcomes you enter, whether that's your own forecasts, a set of Polymarket or Kalshi prices you looked up yourself, or a hypothetical scenario you want to test. Your result is plotted directly against the four real, sourced benchmarks from this piece: random guessing, converted superforecaster performance, and both platforms' published aggregate figures.
DN Proprietary Instrument
DN Prediction Market Calibration Scanner
Enter any set of predicted probabilities and real outcomes to compute a real Brier score, benchmarked against Polymarket, Kalshi, superforecasters, and random chance.
Load an example pattern
Brier score = mean[ (predicted probability − actual outcome)² ], where predicted probability is expressed as a decimal (0 to 1) and actual outcome is 1 (Yes happened) or 0 (No happened), averaged across all events entered. 0 is a perfect score; 0.25 is what always guessing 50/50 produces on a binary event; scores above 0.25 indicate a forecaster who is worse than a coin flip, typically from confident, wrong calls.
Benchmarks shown use the standard 0-to-1 scale: random guessing (0.25, definitional), Polymarket's aggregate resolved-market Brier score (~0.084, per Polymarket's own cited Dune dashboard analysis), Kalshi's Brier score at market close (~0.02, from a 2.24-million-market study), and converted superforecaster performance (~0.05-0.07, converted from the Good Judgment Project's original doubled-scale published figures of below 0.12). The conversion matters: some superforecasting literature reports Brier scores on Tetlock's original 0-to-2 scale, where guessing 50/50 yields 0.5 rather than 0.25, and comparing those figures directly to modern market studies without converting first is a common, significant error.
DN Prediction Market Calibration Scanner is an illustrative educational model, not financial advice. Brier score benchmarks are drawn from published third-party research and are subject to each study's own methodology and sample. Not a recommendation regarding any specific market or platform. May be reproduced with attribution to decentralised.news.
What would actually settle this
Two things would move this from "genuinely mixed" to a clearer verdict. If a future high-profile, contested event, another close election, a major geopolitical binary, produces a repeat of 2024's pattern, the highest-volume, most widely cited platform underperforming smaller, position-capped rivals on hit rate specifically, that would be strong evidence the whale-driven distortion explanation is structural rather than a one-off. If instead Polymarket's hit rate on the next comparably scrutinized event matches or beats Kalshi and PredictIt, that would suggest 2024 was an anomaly rather than a predictable weakness. Separately, watch whether platforms and their advocates start reporting hit rate on marquee events specifically, rather than defaulting to the more flattering aggregate calibration figure, which is a different and less contested claim than the one usually implied.
Positioning around a genuinely mixed picture
None of this is a signal to trust or distrust either platform broadly; the calibration research is genuinely strong at scale, and the hit-rate weakness is specific to a narrower, if high-profile, sample. For readers looking to access these platforms directly and form their own view market by market, Bybit and OKX provide the broader crypto infrastructure many prediction-market users rely on for related trading and settlement.
Frequently asked questions
DN-internal: This piece connects to the DN Super Cycle Confirmation Gauge's named-analyst-sensitivity framework and the DN Sentiment Index backlog outline, both testing how much a headline verdict depends on which specific data source or metric is used to produce it.
Sources: DL News, "Are Polymarket and Kalshi as reliable as they say? Not quite, study warns"; Yellow, "Kalshi Study Finds Prediction Markets Work Best When Traders Crowd In" (Kagan & Baiocchi working paper); Kalshi Research, "Prediction market calibration: How accurate is Kalshi?" (August 2026); FrenFlow, "How Accurate Is Polymarket? What the Data Actually Shows" (citing Dune analysis by Alex McCullough and a December 2025 Vanderbilt University study); TradeTheOutcome, "Polymarket Accuracy Report Data 2026"; arXiv, "Price as Focal Point: Prediction Markets, Conditional Reflexivity, and the Politics of Common Knowledge"; EmergentMind, "Superforecasters: Metrics and Methods"; Good Judgment, "The Superforecasters' Track Record"; AI Impacts, "Evidence on good forecasting practices from the Good Judgment Project."
As of: September 10, 2026. Not financial advice. This is high-risk, YMYL content covering contested forecasting-accuracy research; figures reflect the most recent verified reporting available at time of writing and may be revised as underlying studies are updated. The Prediction Market Calibration Scanner is an illustrative educational model, not a live feed.






