Your forecasts
Illustrative scenario · no live market data
Sample evaluation
Brier = mean((probability − observed result)²).
Lower is better. The reference assigns 50% to every event and scores 0.25. A small illustrative sample cannot establish future accuracy or trading profitability.
What the score is and how the calculator works
For each question you enter a probability between 0 and 1 and the outcome after resolution, 1 if it happened and 0 if it did not. The score is the mean of (p - y)^2 across all rows, so it ranges from 0 to 1 on this unscaled binary convention and lower is better. The outcome field accepts 0 or 1 after resolution.
The Brier score is a proper scoring rule, studied in the literature as a way to evaluate probabilistic forecasts. A 2008 paper by Jochen Bröcker on reliability and the decomposition of proper scores discusses binary forecasts and the Brier score in its abstract, but it provides no empirical prediction-market or trading results.
Worked examples
If you forecast 0.8 and the outcome is 1, the squared error is (0.8 - 1)^2 = 0.04. If you forecast 0.8 and the outcome is 0, the squared error is (0.8 - 0)^2 = 0.64. Averaging across questions combines these terms, so one confident miss can move your score more than several near-correct calls. To relate probabilities to contract payoffs and costs, use the expected value calculator.
A constant forecast of 0.5 scores (0.5 - y)^2 = 0.25 for every binary outcome, because either y = 0 or y = 1 gives the same squared distance. That constant-0.5 value is a reference point, not an optimal or universal performance threshold, and it should not be confused with the 0-to-2 two-class convention that doubles every score.
Freeze the record before comparing methods
For a meaningful comparison, freeze the question, resolution rule, forecast timestamp, model version, prompts and the sources that were available at the time. Exclude information that arrived after the forecast, record forecasts and observed outcomes, and compare models and baselines on the same questions and horizon. Disclose any exclusions and missing predictions rather than dropping them silently, and calculate the binary Brier score separately from calibration, sample size, uncertainty and costs. This is a proposed PolyZeno workflow; it has not been run as an experiment, and no model ranking or profitability result is claimed.
One mixed sample of resolved questions limits what you can infer. A lower average on those rows does not prove the method is calibrated, that it will score well on a new sample, or that it would be profitable to trade. Small samples make the average sensitive to a few rows, and different question selections or time horizons change the comparison. To connect forecast quality with contract payoffs, read about expected value and evaluating AI forecasts.
Sources & verification
PolyZeno. Automated review with DeepSeek V4.1 Flash.