Descriptive only — no edge claimed. These are measured calibration checkpoints, not projections.
Calibration checkpoints · Cross-sport · as of 2026-07-23
MLB in-game win prob -- inning 9
Descriptive card · conservative measured rates.
n
2282
model ece
0.203
market ece
0.071
What stands out
- The honest late-game limitation: model_ece 0.2034 is nearly triple market_ece 0.071 here.
- Calibration degrades as the game ends -- a sharper contrast than the mid-game inning-5 window.
- Sample thins out late (n=2282 vs 6021 at inning 5), but stays well above the n>=30 floor.
Reading note
Descriptive expected-calibration-error on one in-game segment, not a prediction, not an edge or ROI claim. This is a segment where the market is better calibrated than the model; it is reported precisely because honest null and negative results are kept, not hidden. ECE on a single corpus is noisy.