Skip to content

Measurement / Calibration

Inspect every reliability bin.

A reliability bin groups forecasts within a published probability range, then compares their mean forecast with the observed frequency. The diagonal is the calibrated reference; a point away from it shows the direction and size of that bin's published gap.

Aggregate reliability bins are a different object from state-conditioned reliability: these bins group forecasts by probability range, while the state inspector also groups by game phase. See state-conditioned reliability for that separate published grid.

Observed-frequency intervals use 1,000 bootstrap resamples clustered by game_id, with a 2.5% to 97.5% interval. Bins use a minimum floor of 5 games; n counts ticks, not games.

Reading trail

Read first: Read the reliability finding before comparing published bins.

Next question: Where do repeated observations reduce the distinct support behind a bin?

Read the analysis: Does calibration hold up over time and across market types? (sources regenerated)

Three different properties

Brier score
The mean squared error of the forecast probability against the 0/1 outcome. Lower is better; it combines calibration and resolution.
Calibration error
ECE and the reliability curve measure how far observed frequency sits from the forecast in each bin. This page shows the published reliability bins and ECE summaries.
Sharpness
How far forecasts sit from the base rate. Sharpness is different from Brier and calibration error, and this page does not show a sharpness measure.

Reliability curves with uncertainty

Exact fields: calibration_stability.json -> sports[sport].sides[side].bins[] -> mean_p, mean_y, mean_y_ci, gap, gap_ci, n, n_games, and low_n.

MLB: 27,351 ticks from 178 games; ticks are not independent games.

Perfect calibrationModelMarketLow-n bin
0%0%25%25%50%50%75%75%100%100%Mean forecast probabilityObserved frequencyModel 5.1% forecast, 6.3% observed, 1.19 pp gap, 2,072 ticksModel 15.6% forecast, 4.6% observed, -10.99 pp gap, 1,900 ticksModel 24.5% forecast, 10.7% observed, -13.84 pp gap, 1,774 ticksModel 35.8% forecast, 23.5% observed, -12.32 pp gap, 2,217 ticksModel 45.8% forecast, 37.1% observed, -8.67 pp gap, 3,859 ticksModel 54.3% forecast, 55.2% observed, 0.87 pp gap, 5,636 ticksModel 64.7% forecast, 64.3% observed, -0.43 pp gap, 4,026 ticksModel 74.6% forecast, 82.2% observed, 7.52 pp gap, 2,024 ticksModel 84.9% forecast, 94.1% observed, 9.27 pp gap, 2,406 ticksModel 95.2% forecast, 98.0% observed, 2.82 pp gap, 1,437 ticksMarket 4.9% forecast, 2.1% observed, -2.78 pp gap, 3,233 ticksMarket 14.8% forecast, 12.1% observed, -2.69 pp gap, 2,344 ticksMarket 24.8% forecast, 15.9% observed, -8.85 pp gap, 1,876 ticksMarket 35.1% forecast, 23.7% observed, -11.43 pp gap, 1,526 ticksMarket 45.3% forecast, 34.5% observed, -10.75 pp gap, 2,703 ticksMarket 55.0% forecast, 49.3% observed, -5.66 pp gap, 4,129 ticksMarket 64.9% forecast, 68.2% observed, 3.31 pp gap, 2,957 ticksMarket 74.6% forecast, 70.4% observed, -4.24 pp gap, 2,579 ticksMarket 85.2% forecast, 86.3% observed, 1.10 pp gap, 2,449 ticksMarket 94.9% forecast, 97.5% observed, 2.53 pp gap, 3,555 ticks

MLB / Model worked example

0% to 10% probability bin

Mean forecast 5.1% and observed frequency 6.3% produce a published 1.19 pp gap. The observed-frequency interval is 1.4% to 13.4%.

Eligible bins
10
Significant bins
3
Within noise
7
Published Brier
0.166759

MLB / Market worked example

0% to 10% probability bin

Mean forecast 4.9% and observed frequency 2.1% produce a published -2.78 pp gap. The observed-frequency interval is 0.0% to 5.6%.

Eligible bins
10
Significant bins
0
Within noise
10
Published Brier
0.151008
Published reliability bins for MLB
SeriesBin rangeMean forecastObservedGap (pp)Gap CI (pp)n ticksn gamesLow n
Model0% to 10%5.1%6.3%1.19 pp-3.66 pp to 8.21 pp2,07286No
Model10% to 20%15.6%4.6%-10.99 pp-15.30 pp to -4.56 pp1,90073No
Model20% to 30%24.5%10.7%-13.84 pp-21.57 pp to -4.48 pp1,77469No
Model30% to 40%35.8%23.5%-12.32 pp-23.03 pp to 1.39 pp2,21780No
Model40% to 50%45.8%37.1%-8.67 pp-21.95 pp to 5.12 pp3,859100No
Model50% to 60%54.3%55.2%0.87 pp-10.26 pp to 12.17 pp5,636114No
Model60% to 70%64.7%64.3%-0.43 pp-14.75 pp to 13.40 pp4,02695No
Model70% to 80%74.6%82.2%7.52 pp-3.76 pp to 17.75 pp2,02470No
Model80% to 90%84.9%94.1%9.27 pp2.98 pp to 13.76 pp2,40681No
Model90% to 100%95.2%98.0%2.82 pp-0.63 pp to 4.88 pp1,43761No
Market0% to 10%4.9%2.1%-2.78 pp-4.84 pp to 0.55 pp3,23383No
Market10% to 20%14.8%12.1%-2.69 pp-11.03 pp to 9.43 pp2,34487No
Market20% to 30%24.8%15.9%-8.85 pp-17.20 pp to 2.90 pp1,87690No
Market30% to 40%35.1%23.7%-11.43 pp-22.02 pp to 0.24 pp1,52698No
Market40% to 50%45.3%34.5%-10.75 pp-23.20 pp to 3.14 pp2,703127No
Market50% to 60%55.0%49.3%-5.66 pp-16.18 pp to 5.95 pp4,129140No
Market60% to 70%64.9%68.2%3.31 pp-8.70 pp to 13.75 pp2,957117No
Market70% to 80%74.6%70.4%-4.24 pp-18.01 pp to 8.23 pp2,579100No
Market80% to 90%85.2%86.3%1.10 pp-8.35 pp to 9.11 pp2,44988No
Market90% to 100%94.9%97.5%2.53 pp-0.63 pp to 4.48 pp3,55583No

These reliability curves cover MLB and international soccer only. See calibration by game checkpoint for checkpoint evidence from other published corpora.

Summary by market

Published Brier and expected calibration error scores, where resolved outcomes are available. Lower Brier means lower average squared probability error.

SportMarketModel BrierMarket BrierModel ECEMarket ECEnPublished comparison
MLBMLB in-game moneyline (game winner)0.1667590.1510080.0573670.04905727,351Market has lower Brier
International soccersoccer in-game match result (2-way, as graded)0.3325220.1951570.4002360.2818664,265Market has lower Brier
MLBMLB in-game over/under run total (Kalshi KXMLBTOTAL quote-tracks)Not scoredNot scoredNot scoredNot scored2,113No resolved outcome label

The scored MLB and international soccer rows publish lower market Brier values. The total row remains unscored because its source has no resolved outcome label.