Lab
Instruments & counterfactuals
Descriptive instruments and counterfactual replays. edge_claimed:false throughout -- these surfaces measure, they do not claim.
Counterfactual explorer
Pick a game-state cell -- probability band x game-time -- and see how well the model was calibrated there, next to the market. Reads committed cells; runs no new simulation.
MEASURE calibration error only|edge_claimed: false on this artifactState-conditioned calibration grid (MLB)
Model calibration error per (band x game-time) cell -- click a cell
late x .8-1
| source | n | mean p | obs freq | cal err |
|---|---|---|---|---|
| MODEL | 5,037 | 0.9211 | 0.6853 | 0.2357 |
| MARKET | 2,771 | 0.9199 | 0.9405 | 0.0206 |
In this cell the model's mean forecast was 0.9211 and outcomes landed at 0.6853; calibration error 0.2357. Lower is better; the market's error here was 0.0206.
Worst buckets = the improvement backlog
Ranked-worst model cells -- shown as a feature, not hidden
| cell | model n | model cal err | market cal err | read |
|---|---|---|---|---|
| late((inn7+)) x .8-1 | 5,037 | 0.2357 | 0.0206 | MODEL TRAILS |
| late((inn7+)) x 0-.2 | 5,431 | 0.1528 | 0.0336 | MODEL TRAILS |
| early((inn1-3)) x .6-.8 | 5,098 | 0.1145 | 0.0325 | MODEL TRAILS |
| mid((inn4-6)) x .8-1 | 2,718 | 0.0837 | 0.0076 | MODEL TRAILS |
Every figure traces to a committed JSON under scripts/platformkit/analytics_showcase/out/. Grid cells read verbatim from state_conditioned_calibration.json. No ROI, profit, or betting-edge claim appears on this page -- calibration error only.