Forecast calibration: compare probabilities with outcomes

Calibration tests whether your 70% forecasts resolve YES about 70% of the time. Worked bucket, Brier arithmetic, and why small samples mislead.

In this guide

What calibration means, in plain terms

Calibration is the match between the probability you attach to an event and how often events like it actually happen. Say 70% 100 times across independent questions, and a perfectly calibrated forecaster expects roughly 70 YES resolutions. Being calibrated is not the same as being right on any single question, and it is not the same as sharpness, which is how far your probabilities sit from 50%. Someone who always writes 50% can look well calibrated on balanced questions while telling you almost nothing.

On a platform that aggregates probabilistic forecasts into community predictions instead of selling event shares, the underlying object is still a probability attached to a question with objective resolution criteria. Prize structures vary between tournaments, and question series do not automatically carry rewards, so you cannot assume one universal scoring or payout rule. Calibration is a diagnostic on your own forecasts, kept separate from any payout.

A worked bucket: 10 70% predictions

Take an illustrative bucket of 10 independent questions where you wrote 70% each time. 7 of them resolve YES. The mean forecast, the average of your 10 probabilities, is 0.7. The observed frequency, the share of YES outcomes, is also 0.7. That bucket is calibrated in the simplest possible sense: the bin's average probability equals its resolution rate.

This is illustrative arithmetic, not an empirical study. The numbers are easy to follow on purpose, not because 10 forecasts can settle anything durable about your skill.

Brier arithmetic inside the same bucket

The Brier score is a mean squared error on probabilities: for each question, subtract the outcome coded as 1 for YES or 0 for NO from your probability, square it, then average across the bucket. 7 YES resolutions at 70% contribute 7 times (0.7 - 1)^2 = 7 times 0.09 = 0.63. 3 NO resolutions at 70% contribute 3 times (0.7 - 0)^2 = 3 times 0.49 = 1.47. Total: 2.10 across 10 questions, so the mean Brier score is 0.210.

Now add one settled question, forecast at 20% and resolved NO. Its squared error is (0.2 - 0)^2 = 0.04. Total becomes 2.10 + 0.04 = 2.14 across 11 questions, mean 0.1945. The isolated figure of 0.065 from a smaller mixed set would correspond to a different bucket composition entirely, and the gap shows how much bucket membership moves the number. What matters for calibration is not one Brier value but whether errors line up systematically with your probability bins: high probabilities that keep resolving NO, or low probabilities that keep resolving YES, reveal miscalibration.

Calibration, sharpness and profitability are 3 different questions

A forecaster can be calibrated and useless: always writing 50% on balanced binary questions yields decent calibration and zero sharpness. A forecaster can be sharp and miscalibrated: writing 99% often and being right 80% of the time is confident but overconfident. A useful review looks at both dimensions, bin by bin, with counts attached to each bin.

Calibration does not imply profitability. A well-calibrated probability still loses money if the market price already reflects it, if fees and spreads eat the edge, if position sizing is wrong, or if the event resolves on a technicality. The losing case is the ordinary one: your 70% is honest, the contract trades at 72 cents, and you pay 72 to earn 70 cents of expected value. Nothing in a calibration bucket tells you a trade would have made money.

A reading aid for one bucket

The table below is a decision aid for interpreting a single calibration bucket, not a rule set derived from empirical data. Use it to decide what to record next, not to declare yourself calibrated or uncalibrated.

Illustrative reading aid for a single calibration bucket, not empirical guidance
Bucket patternPlausible readingNext recording step
Mean probability equals YES rateConsistent with calibrationIncrease count in this bin
Mean probability above YES rateOverconfident in this binLog more questions in the bin
Mean probability below YES rateUnderconfident in this binCheck whether probabilities were hedged
Very few entries in binChance match possibleAvoid conclusions until the sample is larger
Entries share a driverDependence weakens sampleSeparate dependent questions

Sample size, selection and dependence

A small bucket can match by chance. 10 forecasts whose mean probability equals their resolution rate look reassuring, but random noise produces the same picture easily at that size. I will not hand you a minimum sample-size threshold, because the honest answer depends on the width of your bins and how much precision you need; overlapping intervals tell you more than a point estimate ever will.

Selection matters just as much. If you only log questions you feel confident about, or only ones that already resolved, your bucket is not a sample of your forecasting behaviour. Write the inclusion rule before you start: which questions enter the journal, when you record the probability, and what counts as resolution.

Independence is another assumption to check. 10 forecasts about the same event or the same underlying driver are not 10 independent draws, and counting them as such inflates the apparent sample size without adding information.

A recording protocol you can reuse

Each journal entry needs the question text, the resolution criterion, the probability you assigned, the timestamp, and the final outcome. When you review, group entries into bins such as 0 to 20%, 20 to 40%, and upward, then compare the mean probability per bin with the observed YES rate. Report counts next to percentages so a reader can see how thin each bin is.

A journal tool is planned but not yet available; treat it as unshipped until a link to it exists. Until then a plain spreadsheet with those columns runs every check described here.

Where to go next

For a deeper numeric treatment of the squared-error metric used above, read brier score explained. If you are weighing machine estimates against your own, start with AI forecasting. And for the difference between a probability and a statistical expectation, which trips up many journal reviews, expectation versus probability lays out the distinction.

Sources & verification

Metaculus: FAQ ↗

Sources checked

PolyZeno. Automated review with DeepSeek V4.1 Flash.