Backtest a prediction-market strategy without future leakage
Split history by decision time, keep closed markets and failed entries, and record executable quotes, depth and fees instead of midpoints.
In this guide
What future leakage looks like in a prediction market
A prediction-market backtest asks a narrow question: if I had followed these rules on that day, would I have ended up ahead or behind? You can answer it only with information that existed on that day. Any material that appeared after resolution, from final odds to a summary written weeks later, turns the test into a description of what already happened.
Say you record the midpoint of a book as your entry price. The midpoint is not a price anyone offered you in size; a real order walks the book and eats whatever depth is stacked at each level. So write executable quote when you mean the price and size you could actually have taken, with the side and the depth you needed. Displayed prices on Polymarket are usually the midpoint of the best bid and ask, and when the spread is wider than 0.10 they switch to the last trade, which means neither number is a guaranteed fill.
Split by decision time, not by row
The protocol here is my own design and illustrative. It has not been run against a market dataset, and no benchmark numbers back it. Sort markets by the timestamp of the decision, then cut that timeline into 3 consecutive blocks: train on the oldest, validate in the middle, test on the newest. Any ratio you pick is a starting heuristic, not a rule from the sources; what matters more than the ratio is the order. Shuffling market events across blocks puts a later market in an earlier block and hands your model information it would not have had.
At each decision point the ledger should hold what you could see then: the executable quote, the depth available at that quote, the fees you knew about that day, and the market status. Then keep the ugly rows. A market that closed before your signal, a market with no depth at your limit, a market where your order never filled: those are results, not blanks. Deleting them is the cheapest way to make a strategy look profitable.
If you are scoring forecasts rather than trading shares, say so plainly, because the 2 are not the same object. Metaculus users submit probabilistic forecasts that are aggregated into community predictions, and questions resolve against objective criteria; some tournaments carry performance prize pools while question series do not necessarily have prizes. A backtest of forecast accuracy is not automatically a backtest of tradable profit, and the FAQ does not promise a single scoring implementation or universal cash rewards.
Worked example: from gross to net, and the losses you forgot
Start with a hypothetical winning trade. Gross profit before any cost is 20 units. Entry and exit costs, counting fees and the spread you crossed, add to 8 units. The net is 20 minus 8, or 12 units. That arithmetic is fine as far as it goes, and it is still not evidence that the strategy works.
The supplied protocol adds a silent test: if you omitted losses from this ledger, the 12 is wrong and the omitted losses invalidate the performance. The supplied example does not state how many losses there were or how large they were, and I will not invent them. The honest next step is to open the full ledger, count every failed or closed entry, and recompute. All figures here are illustrative, not measured.
Treat this as a template for your own columns, not as a data point. Spread width, depth and slippage decide whether the number you write in the entry-price cell resembles what your order would have paid in a live book.
| Line | Input or outcome | Value |
|---|---|---|
| Winning entry | Gross profit before costs | 20 |
| Winning entry | Entry and exit costs, fees and spread | 8 |
| Winning entry | Net result of the reported trade | 12 |
| Failed or closed entries left out | Not supplied in the example | Unknown |
| Full ledger | Effect of omitted losses on reported performance | Invalidates performance until recomputed |
When a language model appears to forecast the past
A model trained before a market resolved may already have seen the resolution in its training text. Ask it later to forecast the same market and it can reproduce the answer while sounding like analysis. This is future leakage wearing a different coat: it arrives as a confident sentence rather than a suspicious column, which is exactly why it is easy to miss.
A prospective record is the honest way to check whether a model can forecast at all. Before the market resolves, save the forecast, the timestamp, the exact prompt and everything the model could read at that moment. After resolution, compare the saved forecast with the outcome. A prompt about a historical event, however carefully worded, cannot substantiate an unseen prediction, because the model may simply know how the story ended. No benchmark results have been produced for this protocol.
Limits you should keep in view
A backtest cannot promise future profit. It reports how rules behaved under recorded assumptions, and those assumptions can be wrong in ways that matter: quotes go stale, depth thins out right when your signal fires, fee schedules change, and sometimes nobody is on the other side. A chronological split removes one class of leakage and does nothing about selection bias, survivorship bias, or shifts in how the venue operates.
Data-source mismatch is the other common trap. Forecast platforms produce probabilities and scores; a trading venue produces quotes and fills. The Metaculus FAQ makes that distinction clear, describing community predictions and objective resolution criteria while noting that prize pools exist in some tournaments and not in question series generally. Mixing a scoring series with a quote series can make a strategy look tradable when it never was.
Your next move this week
If you already have a price ledger, add the 2 columns that are usually missing: market status and failed entry. Re-run the rules with those rows present. A result that changes sign under that change has taught you more than a smooth equity curve would. Then open a prospective folder and save the next language-model forecast before its market resolves, prompt and timestamp included.
The AI overview explains how to evaluate forecasts. The AI workflow guide turns dated evidence into a repeatable analysis process, while the liquidity guide helps you check assumptions about fills.
To measure a model’s prediction errors, see evaluating AI forecasts.
Sources & verification
Sources checked
Sources checked
PolyZeno. Automated review with DeepSeek V4.1 Flash.