Browse › Calibration Stability
Analytics module · as of 2026-07-23
Bootstrap by game: the market's MLB calibration is statistically perfect
mlb/model: 4/10 eligible bins have a calibration-gap 95% CI that excludes 0; 6 within noise. mlb/market: 0/10 eligible bins have a calibration-gap 95% CI that excludes 0; 10 within noise. soccer_intl/model: 8/10 eligible bins have a cali...
confirmednull (a finding)not testabledescriptivepending
Calibration Stability

What it means
A cluster bootstrap (resampling game_id, 1000 draws) asks which calibration gaps survive sampling noise. For the market, none do -- its in-game MLB prices are indistinguishable from perfectly calibrated. The model has four real misses, the largest being overconfidence on near-locks: the 0.9-1.0 bin predicts 0.9541 but wins only 0.6632.
Caveats & confounds
Soccer is much noisier -- 8 of 10 model bins and even 6 of 10 market bins deviate significantly on just 51 games, so the small corpus limits any firm conclusion there.
Method. cluster bootstrap (resample game_id with replacement, re-bin) on 10-bin reliability curves
Ask Scout about this