How to evaluate an AI forecasting model
To evaluate an AI forecasting model, you compare its probability forecasts against resolved outcomes on questions and horizons that were fixed before the outcomes were known.
In this guide
Why a frozen evaluation matters
PolyZeno proposes a reproducible evaluation protocol: freeze question, resolution rule, forecast timestamp, model version, prompts and then-available sources; exclude later information; record forecasts and observed outcomes; compare models and baselines on the same questions and horizon; disclose exclusions and missing predictions. Calculate binary Brier and examine calibration, sample size, uncertainty and costs separately. No model ranking or experiment results have been measured for this site. This is a proposed workflow, not evidence that any model has been tested or that any strategy is profitable.
The protocol is built around the idea that a prediction only means something relative to what was knowable when it was made. If you let the model see the answer, or if you grade it on questions you selected after the fact, the evaluation measures hindsight rather than forecasting ability. A confident explanation alone cannot establish accuracy, so the protocol insists on recording the inputs and the timestamp before resolution.
Define the inputs and the test set
A reproducible evaluation starts by freezing the question, the resolution rule, the forecast timestamp, the model version, the prompts and the sources available at that time. Then you exclude any information that emerged later, including the resolved answer and any reporting about it. Record the actual forecast and the observed outcome for every question, and compare models and baselines on the same questions and the same horizon. If a model or baseline has no prediction for a question, that is a missing prediction, and you should disclose it rather than delete the question. Comparisons are only meaningful when the set of questions is identical and the selection rule for that set is stated.
The proposed protocol does not prescribe a specific model or baseline, but it does require that whatever you compare is compared on the same questions and horizon. A model that looks better on a hand-picked set of easy questions may look worse on a representative set, so the selection rule matters as much as the score. Missing predictions are part of the result: if a model skips hard questions, you cannot claim it is accurate on those questions.
This error covers 1 forecast. Model accuracy requires a sample of resolved outcomes.
Use Brier score and go beyond it
For binary questions, the Brier score is the mean of squared differences between the forecast probability p and the resolved outcome y, where p is in [0,1] and y is either 0 or 1. In the unscaled binary convention, the score ranges from 0 to 1 and lower is better. A forecast of .8 scores .04 when the outcome is 1, and .64 when the outcome is 0. A constant forecast of .5 scores .25 for every binary outcome. Use our <a href='/tools/brier'>Brier score calculator</a> to compute this on your own resolved set, but remember that this single score does not establish calibration, skill on another sample, or trading profitability. Do not call .25 an optimal or universal threshold, and do not confuse it with the 0-to-2 two-class convention.
A single Brier number hides structure. Research on proper scoring rules decomposes probabilistic forecasts into reliability and resolution components, which helps separate a model that is well-calibrated from one that merely makes sharp distinctions. The proposed protocol asks you to examine calibration, sample size, uncertainty and costs separately. With a small sample, the Brier score is noisy, so report how many questions were resolved and how they were selected. Costs matter too: a forecast can be accurate and still lose money after fees, spread and slippage, so evaluate those conditions separately rather than folding them into the probability score. See <a href='/concepts/expected-value'>expected value</a> for the arithmetic of costs and payoffs, and the <a href='/editorial-policy'>editorial policy</a> for how this site treats evidence and uncertainty.
Sources & verification
PolyZeno. Automated review with DeepSeek V4.1 Flash.