Match result accuracy
How often what we forecast on match result matches what actually happened, published unedited — the same numbers the overview shows, just this market alone.
In one sentence
Across 85,308 held-out matches, we are well calibrated on match result, an average gap of -0.0pp between what we said and what happened. The chart below shows each probability band against how often it actually came in.
85,308 matches · model 20260904-0011
What we said, against what happened
Each dot is a tenth of the probability range: what we said, against how often it happened. On the dashed line we were exactly right; above it we were too cautious, below it too confident. The rule through each dot is one standard error, and the dot's area is how many matches it rests on — the sparse ones at the ends will always scatter.
Furthest out: when we said 92%, it happened 100% of the time, across 22 matches.
All 12 gate values for this market, unedited
| Gate | Value | Result |
|---|---|---|
| ece home | 0.0065 | pass |
| max reliability gap pp home | 7.8184 | fail |
| max reliability gap pp home measurable | 2.3105 | pass |
| reliability within sampling envelope home | 2.3105 | pass |
| ece draw | 0.0055 | pass |
| max reliability gap pp draw | 9.7752 | fail |
| max reliability gap pp draw measurable | 2.5280 | pass |
| reliability within sampling envelope draw | 2.5280 | pass |
| ece away | 0.0094 | pass |
| max reliability gap pp away | 6.2980 | fail |
| max reliability gap pp away measurable | 4.2311 | fail |
| reliability within sampling envelope away | 4.2311 | pass |
Common questions
How accurate are your match-result predictions?
We publish how often the outcome we rated most likely actually happened, decile by decile. The chart shows what we said against what happened, and the summary above states in one sentence whether we were well calibrated, slightly over-confident, or slightly under-confident on this market across the held-out matches we checked.