Browse › State Conditioned Calibration
Analytics module · as of
The model's biggest miss: overconfident late-game favorites
Late in MLB games, when the model calls a team 0.9211 to win, it actually wins only 0.6853 of the time -- a 0.2357 calibration error and the model's single biggest well-sampled miss (n=5037).
confirmednull (a finding)not testabledescriptivepending
No committed chart for this module — the receipts below are the evidence.
What it means
This exhibit buckets every graded in-game prediction by confidence band and game phase to find where the model drifts furthest from reality. The worst high-sample bucket is late-inning heavy favorites: the model is overconfident, treating near-locks as more certain than they turn out to be. Tinier buckets look even worse but rest on almost no data.
Caveats & confounds
The very worst rows in ranked_worst_buckets have tiny n (the top one, calibration_error 0.7883, is only n=10) and are noise-dominated; the n-weighted ECE summary (mlb model 0.079 vs market 0.0591) is the stable read.
Ask Scout about this