Explainer
What Calibration Means, and Why We Grade Ourselves Against the Market
Being accurate is the easy bar. Matching a sharp closing line within noise is the real one -- and here is exactly where we fall short.
A calibrated forecast is one whose probabilities mean what they say: of all the times it calls something 70 percent, that thing should happen about 70 percent of the time. The scorekeeping tool is the Brier score — the mean squared error of a probability — where lower is better and 0.25 is a coin flip on a balanced outcome. Accuracy alone is cheap. The demanding test is whether the forecast is as sharp as the betting market's own devigged closing price, because that price is the sharpest public estimate that exists.
Pregame, on team-strength markets, the honest answer is that the model matches the close and beats nothing. On NBA moneyline the model's Brier is 0.1735 against the devigged close's 0.1672 over 372 games — a gap of +0.0063, labeled MATCH within sampling noise. On MLB moneyline it is 0.2429 against 0.2390 over 13,992 games, gap +0.0039, also MATCH. Tracking a sharp close within noise is the ceiling for an efficient market, and hitting the ceiling is the result, not a shortfall.
In-game, the market is simply better, and rather than report that in one deflating sentence we take the gap apart. A Murphy decomposition splits Brier into reliability (miscalibration, fixable by recalibrating) minus resolution (information, which recalibration cannot manufacture). On MLB in-game win probability the model scores 0.237684 against the market's 0.206653 — a gap of +0.0310 — of which the reliability part is only +0.0066 and the resolution part is -0.0235. In plain terms: the market is not ahead because our model is badly calibrated; it is ahead because it resolves more information into its price than our checkpoint snapshots can see. Soccer tells the same story on a smaller, noisier corpus (model 0.227887 versus market 0.142726). This matches the standing finding that an early in-game deficit is a news gap, not a model gap.
We also publish where the model is weakest, as a ranked backlog rather than a hidden weakness. Bucketing every graded in-game prediction by probability band and game-state time, the ugliest MLB cell is late-inning heavy favorites: in innings seven and later, in the 0.8-to-1.0 band, the model predicts 0.9211 but those teams only win 0.6853 of the time over 5,037 rows — badly overconfident, while the market's error in the same band is a fraction of that. The single worst-ranked bucket overall is a late-soccer, low-probability cell with only 10 rows, and it is flagged as top-of-backlog but not statistically load-bearing. Naming the loss, small-sample caveats included, is the point.
There is exactly one measured calibration win, and it is labeled precisely. Fusing the pregame rating prior with the realized mid-game state sharpens the win-probability forecaster: Brier improves from 0.209 to 0.159 in NBA and from 0.241 to 0.126 in MLB on real out-of-sample corpora. It is calibration, not an edge — a live book sees the same score at the same moment, so there is no beat-the-market claim to make. Most of the lift is mechanical (the scoreboard tells you a lot by mid-game); over a score-only baseline the model's own prior adds roughly a quarter of it in NBA and almost nothing in MLB, and we state it that way.
Grading yourself against the sharpest available benchmark, decomposing the gap into fault versus information, and publishing the buckets you handle worst is the discipline. The number that looks good is easy to find; the instrument that locates the loss is the work.
Sources
docs/evidence/calibration-decomposition.mdwebapp/public/data/ask/calibration-market.jsonwebapp/public/data/ask/system-honesty.jsonNext explainer
The Graveyard: Why We Publish What Failed →