Prediction markets vs Metaculus: prices and forecast aggregation
How Polymarket order-book prices differ from Metaculus aggregated forecasts, with a worked 0.70 vs 0.65 case, spread caveats, and a decision rule.
In this guide
2 different objects: an executable quote and an aggregated forecast
On a prediction market such as Polymarket, an event question is a contract you can buy or sell. The number you see on the page is usually the midpoint between the best bid and the best ask, so it is a reference price built from resting orders, not a price anyone has to honour. When the spread between best bid and best ask is wider than 0.10, the platform shows the last traded price instead, which is a past transaction rather than a standing offer. If you send a market order large enough to cross several levels of the book, you pay slippage: the average fill price drifts away from the midpoint, and the midpoint is never an executable guarantee.
Metaculus works differently. You submit a probability for a question that resolves against objective criteria, and the platform aggregates submissions into a community prediction. There is no share buy and sell mechanism, so you cannot exit a forecast before resolution the way you can sell a contract. Some tournaments attach performance prize pools, but a prize pool is not a return from holding event shares, and a question series does not necessarily carry any prize at all. So the 2 systems hand you different things: one gives a tradable position, the other gives an aggregated probability number.
Why 0.70 is not automatically better than 0.65
A market price of 0.65 reflects the current spread, the depth available at each level, and the observation date. Exchange or gas costs, where they apply, are additional transaction costs rather than being contained in the quote. A Metaculus community prediction of 0.70 is an aggregated subjective probability with no settlement attached by default. Subtracting one from the other looks like a 0.05 edge, but the 2 numbers were produced under different rules, different incentives and different populations of participants. Before treating the gap as information, line up the resolution source, the expiry or horizon, and the exact wording of the question. If those do not match, you are comparing 2 different questions, and the difference tells you nothing.
Scores make the confusion worse. A Metaculus score is a performance measure on resolved questions, and the platform does not apply one universal scoring implementation. It is not a cash balance and it does not transfer into financial profit and loss. A good score tells you something about your calibration. It does not tell you that you could have bought at 0.65 and sold at 0.70.
Side-by-side mechanics
The table below compares the 2 systems only on the mechanics documented by the platforms. It says nothing about fees on any particular live question, because those depend on the platform, the market and the observation date.
| Aspect | Prediction market | Metaculus |
|---|---|---|
| What you get | A tradable event contract with a displayed price, usually the midpoint of best bid and best ask | An aggregated community probability from submitted forecasts |
| How you participate | Buy or sell shares through an order book | Submit a probability; no share buy or sell |
| Price caveat | Midpoint is a reference, not executable; spread above 0.10 shows the last trade; large orders incur slippage | No price; the number is a forecast, and aggregation can differ by question |
| Exit before resolution | Possible by selling into the book | Not applicable; you cannot sell a forecast |
| Economic settlement | Contract pays according to market rules | No default cash settlement; resolution follows objective criteria |
| Rewards | Realized profit or loss after entry, exit, costs and resolution | Some tournaments have performance prize pools; series may have none |
| What a score means | Financial profit and loss | Performance measure on resolved questions, not transferable profit and loss |
Worked comparison on one question
Say the same real-world question, with the same resolution source and the same expiry, appears on a prediction market and on Metaculus. The displayed midpoint is 0.65, and your own forecast is 0.70. That looks like a 0.05 difference until you look at the book. Suppose the best bid sits at 0.62 and the best ask at 0.68, an illustrative spread of 0.06. Buying now costs 0.68, not 0.65, so the gap against your 0.70 is 0.02 before any fee, and a size large enough to clear the offer may push the average fill past 0.70 and wipe the gap out. The right response is not to trade on the difference but to log it: question text, timestamp, horizon, market midpoint, executable ask, spread, your probability, and the eventual outcome. After resolution you can compare the realized result with both numbers. That record supports calibration and scoring methods. It does not establish arbitrage, because the market may have priced information you missed, your 0.70 may itself be miscalibrated, or the 2 listings may resolve on slightly different criteria.
| Input | Value | What it illustrates |
|---|---|---|
| Displayed market midpoint | 0.65 | Reference price from the book |
| Best bid | 0.62 | Illustrative only |
| Best ask | 0.68 | What a buy actually pays at best offer |
| Spread | 0.06 | 0.68 minus 0.62; below the 0.10 display threshold |
| Your forecast | 0.70 | Probability you would submit on Metaculus |
| Gap vs midpoint | 0.05 | 0.70 minus 0.65 |
| Gap vs executable ask | 0.02 | 0.70 minus 0.68, before fees and slippage |
| Next step | Record and wait for resolution | Compare realized outcome with both numbers; no arbitrage claim |
When a wide spread kills the comparison
Once the spread passes 0.10, the platform stops showing the midpoint and shows the last trade. At that point neither number is a reliable executable quote, and comparing your forecast against it is mostly comparing against thin or stale liquidity. If the spread is wider than the edge you think you have, there is nothing actionable in the comparison. That is a losing case worth writing down rather than a hidden opportunity worth forcing.
Rewards, scoring and what does not carry over
Some Metaculus tournaments have performance prize pools, some question series may have none, and scoring is not universal across events. What carries over into your market research is the record of forecasts and outcomes, not a monetary balance. If your aim is to compare forecasting skill against market prices over time, a log with entries for the question, timestamp, horizon, your probability, the midpoint, the executable quote when one exists, and the realized outcome is what makes that comparison possible. Keep the claims inside what the log can support. That same log is also what makes AI forecasting approaches comparable when human and model submissions address the same question.
Which one to use for the next question
If you need to hold a position and exit before resolution, use a prediction market and price the spread and slippage into your entry, treating exchange or gas costs as additional where they apply. If you want to state and score a probability without buying shares, use Metaculus. If you want to test whether your forecasts beat market prices, run it as a measurement exercise with recorded timestamps and later outcome checks rather than as a live trade. If you are still deciding which platform's settlement mechanics fit your workflow, the alternatives comparison is the next place to look.
Sources & verification
Sources checked
Sources checked
PolyZeno. Automated review with DeepSeek V4.1 Flash.