Skip to content
BrowseCalibration Stability
Descriptive only — a measured pattern, no edge claimed.
Analytics module · as of 2026-07-23

Bootstrap by game: the market's MLB calibration is statistically perfect

mlb/model: 4/10 eligible bins have a calibration-gap 95% CI that excludes 0; 6 within noise. mlb/market: 0/10 eligible bins have a calibration-gap 95% CI that excludes 0; 10 within noise. soccer_intl/model: 8/10 eligible bins have a cali...
View full size ↗
confirmednull (a finding)not testabledescriptivepending
Calibration Stability
Chart: Calibration Stability -- mlb/model: 4/10 eligible bins have a calibration-gap 95% CI that excludes 0; 6 within noise. mlb/market: 0/10 eligible bins have a calibration-gap 95% CI that excludes 0; 10 within noise. soccer_intl/model: 8/10 eligible bins have a cali...
scripts/platformkit/analytics_showcase/out/calibration_stability.json2026-07-23

What it means

A cluster bootstrap (resampling game_id, 1000 draws) asks which calibration gaps survive sampling noise. For the market, none do -- its in-game MLB prices are indistinguishable from perfectly calibrated. The model has four real misses, the largest being overconfidence on near-locks: the 0.9-1.0 bin predicts 0.9541 but wins only 0.6632.

Caveats & confounds

Soccer is much noisier -- 8 of 10 model bins and even 6 of 10 market bins deviate significantly on just 51 games, so the small corpus limits any firm conclusion there.

Method. cluster bootstrap (resample game_id with replacement, re-bin) on 10-bin reliability curves

Ask Scout about this
Are the predictions actually well calibrated?Is the model overconfident?How big are the corpora behind these calibration numbers?Ask anything →