Where the MLB gap lives: information, not calibration
mlb: model Brier 0.2377 is worse than market 0.2067 (gap=+0.0310); driven mainly by resolution (information, not fixable by recalibration) (reliability_gap=+0.0066, resolution_gap=-0.0235). soccer_intl: model Brier 0.2279 is worse than m...
confirmednull (a finding)not testabledescriptivepending
No committed chart for this module — the receipts below are the evidence.
What it means
A Murphy decomposition splits a Brier score into reliability (are the probabilities honest), resolution (do they separate winners from losers), and an irreducible uncertainty term. The MLB in-game model is almost as well-calibrated as the market (reliability differs by only +0.0066) but separates outcomes less sharply (resolution trails by -0.0235). That means re-scaling the probabilities cannot close the gap -- the market simply knows things the model does not.
Caveats & confounds
Soccer's gap is far larger (model Brier 0.227887 vs market 0.142726) and its reconstructed Brier (0.293529) diverges from the direct Brier, a sign the 10-bin decomposition is coarse on the smaller n=9003 soccer corpus.