Descriptive only — no edge claimed. These are measured calibration checkpoints, not projections.
Calibration checkpoints · Cross-sport · as of 2026-07-23
Soccer (intl) in-game win prob -- minute 0-15
Descriptive card · conservative measured rates.
n
627
model ece
0.306
market ece
0.22
What stands out
- The earliest window is the weakest for both sides -- model_ece 0.3065 and market_ece 0.2204 are both high.
- The model's calibration error runs above the market's here, consistent with an early-game information gap.
- Modest but valid sample: n=627, above the n>=30 render floor.
Reading note
Descriptive expected-calibration-error on one in-game segment, not a prediction, not an edge or ROI claim. Early minutes carry little score signal, so both model and market are poorly calibrated; the model trailing the market here reflects an information gap, not a measured edge in either direction. ECE on a single corpus is noisy at this sample size.