Descriptive only — no edge claimed. These are measured calibration checkpoints, not projections.
Calibration checkpoints · Cross-sport · as of 2026-07-23
MLB in-game win prob -- inning 5
Descriptive card · conservative measured rates.
n
6021
model ece
0.036
market ece
0.049
What stands out
- This is the clean mid-game window: model_ece 0.0363 sits below market_ece 0.0486 on the same checkpoints.
- Sample is substantial for a single inning bucket (n=6021), well above the n>=30 render floor.
- A calibration (reliability) comparison only -- ECE measures how well probabilities match outcomes, not profit.
Reading note
Descriptive expected-calibration-error on one in-game segment, not a prediction, not an edge or ROI claim. ECE on a single corpus is noisy and this is one inning bucket; a lower model ECE here does not generalize to other innings (see inning 9) or to dollars. Market ECE is computed on the same rows for comparison, not as a benchmark the model is claimed to beat overall.