Your forecast history

The journal stays in your browser. Local storage is optional. Export a CSV to keep a copy.

Import replaces the displayed forecasts. Export them first if you want to keep them.

Recorded forecasts0
Resolved forecasts0
Brier score…

Add a forecast or import a CSV.

Calibration by probability range
RangeMean forecastObserved frequencyCount

What you write down, and when

A forecast journal is a log of probabilistic predictions you make before you know the answer. For each question you write the question text, a probability between 0 and 1, and the UTC timestamp of the forecast. When the question resolves, you come back and add the resolved timestamp and an outcome of 1 for yes or 0 for no. Rows with a blank outcome are unresolved, stay in the file, and are excluded from every metric. Scoring an unresolved question would just describe your opinion twice.

The version described here is client-only: the journal is a CSV you keep, there is no network call, no sync and no account. Device-local persistence is optional and only active if you explicitly opt in, with export and clear always available. That suits a log that often contains private reasoning. The cost is that nothing backs itself up, so exporting is the only durable copy you have.

CSV columns and strict validation

The implemented column order is id, question, forecast_at, probability, outcome, resolved_at. Timestamps are UTC in the form YYYY-MM-DDTHH:mm:ssZ. Probability is numeric from 0 to 1. A pending row leaves both outcome and resolved_at blank. A resolved row carries outcome 0 or 1 and a resolved_at strictly after forecast_at. IDs must be unique, no timestamp may be in the future, and the file is capped at 2000 records or 1MB.

Import is atomic: if any row fails, nothing is loaded and the existing file stays untouched. This is stricter than a spreadsheet and it prevents a half-imported journal where some rows are silently discarded. Atomic failure also catches unescaped commas in question text, a probability of 1.5, a duplicate id that would make later matching ambiguous, or a resolution dated before its forecast. Keeping the dataset clean is what makes the metrics below meaningful.

2 rows, then the Brier mean

Brier scoring measures squared error. Take a YES event where you forecast 0.7 and the outcome was 1: (0.7 − 1)² = 0.09. Take a NO event where you forecast 0.2 and the outcome was 0: (0.2 − 0)² = 0.04. The mean across these 2 resolved forecasts is (0.09 + 0.04) / 2 = 0.065. Lower is better, and 0.065 looks strong on this sample, but 2 forecasts say almost nothing about skill. The figure describes those 2 decisions, not your ability.

Illustrative 2-row journal and Brier contributions
idquestionforecast_atprobabilityoutcomeresolved_atbrier contribution
q1Example YES event2026-01-05T12:00:00Z0.712026-02-01T09:00:00Z0.09
q2Example NO event2026-01-06T08:30:00Z0.202026-02-03T14:00:00Z0.04

Calibration buckets and their blind spots

Calibration checks whether the probabilities you state match observed frequencies. Group resolved forecasts into bins, for example 0.6 to 0.7, and compare the bin's mean probability with the share of YES outcomes in it. In the example above, 0.7 and 0.2 fall in different bins, so each bin holds one observation and the comparison is empty of information. Display every nonempty bin with its count rather than hiding thin bins or inventing a minimum cutoff.

A small bucket is a sanity check, not validation and not a profit estimate. It cannot establish statistical reliability, and it says nothing about whether you could profit in a prediction market. Platforms can aggregate probabilistic forecasts into community predictions with objective resolution criteria, and some tournaments carry performance prize pools while others do not, but your journal is only a record. It never implies a reward or an edge, and treating it as one turns it into a vanity dashboard.

Costs, assumptions and the rows that hurt

The direct cost is your time: every entry needs a question, a probability and a timestamp. The main assumption is that outcomes are objective and recorded honestly, because a mislabeled outcome poisons both the Brier mean and the buckets. The losing case is the one worth preserving. A forecast of 0.9 that resolves NO scores 0.81, far worse than the 0.04 above, and a journal that quietly drops those rows is worthless.

A second failure mode is question drift. If you rewrite the question text for an id after resolution, you blur what you actually predicted, so treat forecast_at and probability as immutable in your own process once written, with only resolved_at and outcome appended later. The import checks chronology and uniqueness, not editing discipline, so that discipline is yours. On persistence, optional local storage is convenient but bound to one device; clearing browser data erases it. Export routinely, not only before a cleanup.

Reading the score against outside references

One score on its own is hard to interpret. The Brier score concept gives you the proper background, and pairing it with a calibration guide shows whether your 0.7 forecasts actually resolve YES about 70 percent of the time. An AI-assisted review can help you generate candidate questions or spot vague wording before you commit a probability, but it must not write the resolved_at or outcome fields, because fabricating ground truth destroys the only thing the score depends on.

If you want to test a rule against data that already exists, a backtest workflow is the natural next step, and it needs the same disciplined CSV and the same separation between what you predicted and what actually happened. The journal and the backtest share one principle: the record comes first, the metric second.

Start tomorrow with 3 questions

Create the 6-column CSV, choose 3 questions you genuinely do not know the answer to, and assign each a probability before you look anything up. Use UTC timestamps so timezone drift cannot hide. Import once, confirm validation passes, and stop there. Metrics can wait until something resolves.

After the first resolutions, compute the Brier mean over resolved rows only and print the count beside it. With fewer than 10, label the output illustrative. Do not omit nonempty calibration bins; show every bin count and treat small-sample metrics as illustrative instead of hiding bins. That habit keeps the journal useful after a year rather than misleading within a month.

Sources & verification

Metaculus: FAQ ↗

Sources checked

PolyZeno. Automated review with DeepSeek V4.1 Flash.