Prediction markets vs Metaculus: prices and forecast aggregation

How Polymarket order-book prices differ from Metaculus aggregated forecasts, with a worked 0.70 vs 0.65 case, spread caveats, and a decision rule.

In this guide

2 different objects: an executable quote and an aggregated forecast

On a prediction market such as Polymarket, an event question is a contract you can buy or sell. The number you see on the page is usually the midpoint between the best bid and the best ask, so it is a reference price built from resting orders, not a price anyone has to honour. When the spread between best bid and best ask is wider than 0.10, the platform shows the last traded price instead, which is a past transaction rather than a standing offer. If you send a market order large enough to cross several levels of the book, you pay slippage: the average fill price drifts away from the midpoint, and the midpoint is never an executable guarantee.

Metaculus works differently. You submit a probability for a question that resolves against objective criteria, and the platform aggregates submissions into a community prediction. There is no share buy and sell mechanism, so you cannot exit a forecast before resolution the way you can sell a contract. Some tournaments attach performance prize pools, but a prize pool is not a return from holding event shares, and a question series does not necessarily carry any prize at all. So the 2 systems hand you different things: one gives a tradable position, the other gives an aggregated probability number.

Why 0.70 is not automatically better than 0.65

A market price of 0.65 reflects the current spread, the depth available at each level, and the observation date. Exchange or gas costs, where they apply, are additional transaction costs rather than being contained in the quote. A Metaculus community prediction of 0.70 is an aggregated subjective probability with no settlement attached by default. Subtracting one from the other looks like a 0.05 edge, but the 2 numbers were produced under different rules, different incentives and different populations of participants. Before treating the gap as information, line up the resolution source, the expiry or horizon, and the exact wording of the question. If those do not match, you are comparing 2 different questions, and the difference tells you nothing.

Scores make the confusion worse. A Metaculus score is a performance measure on resolved questions, and the platform does not apply one universal scoring implementation. It is not a cash balance and it does not transfer into financial profit and loss. A good score tells you something about your calibration. It does not tell you that you could have bought at 0.65 and sold at 0.70.

Side-by-side mechanics

The table below compares the 2 systems only on the mechanics documented by the platforms. It says nothing about fees on any particular live question, because those depend on the platform, the market and the observation date.

Documented mechanics: prediction market price vs Metaculus forecast aggregation
AspectPrediction marketMetaculus
What you getA tradable event contract with a displayed price, usually the midpoint of best bid and best askAn aggregated community probability from submitted forecasts
How you participateBuy or sell shares through an order bookSubmit a probability; no share buy or sell
Price caveatMidpoint is a reference, not executable; spread above 0.10 shows the last trade; large orders incur slippageNo price; the number is a forecast, and aggregation can differ by question
Exit before resolutionPossible by selling into the bookNot applicable; you cannot sell a forecast
Economic settlementContract pays according to market rulesNo default cash settlement; resolution follows objective criteria
RewardsRealized profit or loss after entry, exit, costs and resolutionSome tournaments have performance prize pools; series may have none
What a score meansFinancial profit and lossPerformance measure on resolved questions, not transferable profit and loss

Worked comparison on one question

Say the same real-world question, with the same resolution source and the same expiry, appears on a prediction market and on Metaculus. The displayed midpoint is 0.65, and your own forecast is 0.70. That looks like a 0.05 difference until you look at the book. Suppose the best bid sits at 0.62 and the best ask at 0.68, an illustrative spread of 0.06. Buying now costs 0.68, not 0.65, so the gap against your 0.70 is 0.02 before any fee, and a size large enough to clear the offer may push the average fill past 0.70 and wipe the gap out. The right response is not to trade on the difference but to log it: question text, timestamp, horizon, market midpoint, executable ask, spread, your probability, and the eventual outcome. After resolution you can compare the realized result with both numbers. That record supports calibration and scoring methods. It does not establish arbitrage, because the market may have priced information you missed, your 0.70 may itself be miscalibrated, or the 2 listings may resolve on slightly different criteria.

Illustrative single-question comparison (hypothetical inputs, not a tested trade)
InputValueWhat it illustrates
Displayed market midpoint0.65Reference price from the book
Best bid0.62Illustrative only
Best ask0.68What a buy actually pays at best offer
Spread0.060.68 minus 0.62; below the 0.10 display threshold
Your forecast0.70Probability you would submit on Metaculus
Gap vs midpoint0.050.70 minus 0.65
Gap vs executable ask0.020.70 minus 0.68, before fees and slippage
Next stepRecord and wait for resolutionCompare realized outcome with both numbers; no arbitrage claim

When a wide spread kills the comparison

Once the spread passes 0.10, the platform stops showing the midpoint and shows the last trade. At that point neither number is a reliable executable quote, and comparing your forecast against it is mostly comparing against thin or stale liquidity. If the spread is wider than the edge you think you have, there is nothing actionable in the comparison. That is a losing case worth writing down rather than a hidden opportunity worth forcing.

Rewards, scoring and what does not carry over

Some Metaculus tournaments have performance prize pools, some question series may have none, and scoring is not universal across events. What carries over into your market research is the record of forecasts and outcomes, not a monetary balance. If your aim is to compare forecasting skill against market prices over time, a log with entries for the question, timestamp, horizon, your probability, the midpoint, the executable quote when one exists, and the realized outcome is what makes that comparison possible. Keep the claims inside what the log can support. That same log is also what makes AI forecasting approaches comparable when human and model submissions address the same question.

Which one to use for the next question

If you need to hold a position and exit before resolution, use a prediction market and price the spread and slippage into your entry, treating exchange or gas costs as additional where they apply. If you want to state and score a probability without buying shares, use Metaculus. If you want to test whether your forecasts beat market prices, run it as a measurement exercise with recorded timestamps and later outcome checks rather than as a live trade. If you are still deciding which platform's settlement mechanics fit your workflow, the alternatives comparison is the next place to look.

Sources & verification

Polymarket: Prices and order book ↗

Sources checked

Metaculus: FAQ ↗

Sources checked

PolyZeno. Automated review with DeepSeek V4.1 Flash.