AI benchmark markets: score, leaderboard version and cutoff
A threshold market settles on 1 leaderboard tab, 1 control setting and 1 deadline. Check the tab, version and score definition before pricing a contract.
In this guide
The clause that decides a threshold market
A threshold contract does not resolve on the best score that exists anywhere. It resolves on the score shown in a named section of a named leaderboard tab, under a named control setting, by a named deadline. In the captured rule example, the settlement source is the Chatbot Arena LLM Leaderboard and the score comes from the 'Score' section of the 'Text Arena' tab, with style control unchecked, evaluated by December 31, 2026 at 11:59 PM ET.
That leaves 4 things to verify before treating any number as relevant: which tab, which control state, which metric column, and which time condition. A screenshot from a different tab or from a search result is not the settlement source. The same applies when you use a broader model ranking comparison: it is useful context, not the contract source. Ranking rules address who leads a defined list.
Score versus rank, and version matters
A rank is relative: it depends on which other models were on the board at that moment. A score is the value assigned in the metric, so a model can keep a score while its rank changes as other models enter or update. Hugging Face guidance puts it plainly: compare like models and look at the relevant task, because performance on 1 task does not establish performance on another. Academic evaluations and human preference arenas are separate measurements.
Closed API models can change after a dated evaluation, so record the tested version and date. Evaluation contamination can also inflate results by measuring memorized test data rather than capability. If the contract names a leaderboard that updates, the version on the deadline is what counts, not the version you checked earlier.
Worked case: the threshold is 1510
This is an illustrative example from the contract design, not a record of any model. Imagine the required Text Arena score is at least 1510, the deadline is December 31 at 11:59 PM ET, and you see a score of 1512 on a different settings view. You also see 1508 on the required settings: Text Arena, 'Score' section, style control unchecked.
The required-settings score 1508 is below 1510, so the contract condition is not met. The other setting score 1512 comes from a configuration the resolution clause does not name, so it cannot establish YES. The arithmetic is explicit: 1508 >= 1510 is false; 1512 >= 1510 is true, but for the wrong setting. The contract settles on the named setting, so the different-setting 1512 is not evidence for resolution.
| Input | Value | Threshold | Meets contract condition? |
|---|---|---|---|
| Required settings: Text Arena, Score section, style control unchecked | 1508 | >= 1510 | No |
| Other settings view (not named in clause) | 1512 | >= 1510 | No evidence for YES |
Availability: temporary versus permanent
The captured rule example splits source unavailability into 2 cases. If the resolution source is temporarily unavailable, the market remains open until it is accessible again. If it is permanently unavailable, the market resolves to No. That means an outage does not automatically decide the contract, while permanent loss of the source does.
This is a specific clause in a dated example. Do not generalize it to every benchmark market. Read the actual rules for the contract in front of you.
How a binary resolves and what it pays
Ordinary binary winning shares pay 1 and losing shares pay 0; a rare unknown or 50-50 resolution can pay 0.50 per share, not a universal cancellation or refund. The contract’s own rules specify the source, deadline and exceptional cases. Check those clauses for the question being examined, since a title alone cannot determine settlement.
For settlement vocabulary and edge cases, start from the resolution source rather than from the leaderboard screenshot. The same clause that names the leaderboard tab also names the deadline and the unavailability treatment, so read all 3 together before pricing anything.
Forecast skill is a separate measurement
A chat preference or task benchmark score is not a forecast skill measure. Forecast skill needs dated event probabilities and resolved outcomes, scored against what actually happened. Record the event probability before resolution, then preserve the resolved outcome for the chosen evaluation.
A leaderboard score and an event probability answer different questions. For the next step, consult Technology markets, Forecast evaluation and the Brier calculator.
Sources & verification
Sources checked
Sources checked
PolyZeno. Automated review with DeepSeek V4.1 Flash.