Skip to content
The Forecaster

A calibrated engine, measured against itself.

The market is efficient on price — we proved it by rejecting our own pregame signals across three sports. So the honest question is not “can we beat the close” but “does the machinery sharpen the forecast.” It does, in one measured place: mid-game. Every number below wears its receipt, and no dollar edge is claimed.

confirmednull (a finding)not testabledescriptivepending
Walk-forward proof

Trained only on the past, scored on the future.

An expanding-window backtest: every fold asserts max_train_date < min_test_date or fails — no K-fold on time-ordered games. The NBA win-probability ensemble across 3 folds and 2023-24 + 2024-25 seasons; the widening train column is the walk forward itself.

0.709
Accuracy (mean)
+/- 0.025 across folds
0.193
Brier (mean, lower better)
+/- 0.008
3
Expanding folds
2023-24 + 2024-25
73
Leak-checked features
truncation-invariant
Train fracTrain nVal nAccBrier
0.49824910.7370.186
0.614734910.6760.205
0.819644910.7150.188

results/winprob_walk_forward_results.json · 2026-07-20 · edge_claimed: false

The one measured win

In-game conditioning sharpens the forecast.

Fusing the pregame rating prior with the realized mid-game state improves win-probability calibration on a real out-of-sample corpus. A live book also sees the score — so this is calibration, not a claim of beating anyone.

NBA Brier0.2090.159
MLB Brier0.2410.126

Read this as calibration, not a win: most of the drop is the scoreboard state itself, free to anyone watching — the model’s own prior adds only the last sliver, decomposed below.

But how much of that is skill? A rating-blind third arm — conditioning on the score alone, no model prior — splits the lift. Most of it is the scoreboard itself, free to anyone watching. The model’s own contribution is the last column.

Sportstatic (prior only)score-onlycombinedmechanical sharemodel-prior share
NBA0.2090.1720.159~73%-0.014 (~27%)
MLB0.2410.1280.126~99%-0.001 (~1%)

docs/INGAME_PROOF.md Sec. 2 + 2a · real-corpus OOS · each arm rounded independently to 3 dp; the share column is derived from full precision, so it will not reconcile to the 3-dp cells · source doc committed in this repo; live re-run prints VALIDATION_PENDING without the private corpus and falls back to this recorded table · edge_claimed: false

What the model sees

Its own eyes: calibration by game state.

Every graded in-game MLB prediction, bucketed by the model’s probability band and the inning. Each cell is the calibration error — how far the stated probability sits from what actually happened. Darker is a bigger gap: the improvement backlog, in the model’s own view.

MLB in-game calibration error (model)
0-.2.2-.4.4-.6.6-.8.8-1Early (1-3)Early (1-3) / 0-.2: no dataEarly (1-3) / .2-.4: 0.0000.000Early (1-3) / .4-.6: 0.0500.050Early (1-3) / .6-.8: 0.1150.115Early (1-3) / .8-1: no dataMid (4-6)Mid (4-6) / 0-.2: 0.0440.044Mid (4-6) / .2-.4: 0.0270.027Mid (4-6) / .4-.6: 0.0250.025Mid (4-6) / .6-.8: 0.0530.053Mid (4-6) / .8-1: 0.0840.084Late (7+)Late (7+) / 0-.2: 0.1530.153Late (7+) / .2-.4: 0.0330.033Late (7+) / .4-.6: 0.0210.021Late (7+) / .6-.8: 0.0600.060Late (7+) / .8-1: 0.2360.236
scripts/platformkit/analytics_showcase/out/state_conditioned_calibration.json2026-07-23model ECE 0.079 vs market 0.0591
Improvement backlog
Where the forecast is furthest from outcomes
Soccer 75-90+ .2-.4n=10Soccer 75-90+ .2-.4: 0.788 calib. error0.788Soccer 60-75 .2-.4n=90Soccer 60-75 .2-.4: 0.576 calib. error0.576Soccer 30-45 0-.2n=264Soccer 30-45 0-.2: 0.560 calib. error0.560Soccer 75-90+ 0-.2n=379Soccer 75-90+ 0-.2: 0.508 calib. error0.508Soccer 15-30 .4-.6n=120Soccer 15-30 .4-.6: 0.500 calib. error0.500Soccer 0-15 .4-.6n=158Soccer 0-15 .4-.6: 0.494 calib. error0.494
state_conditioned_calibration.json (ranked_worst_buckets)2026-07-23
Cross-sport scoreboard

Model vs market, every checkpoint, reported as-is.

SportMarket @ checkpointnmodel vs market (paired)95% CIverdict
NBAwinprob (Brier) @ end_q11592-0.0084[-0.016, -0.001]market sharper provisional
NBAwinprob (Brier) @ halftime1593-0.0040[-0.010, 0.002]underpowered
NBAwinprob (Brier) @ end_q31593+0.0011[-0.003, 0.005]underpowered
NBAwinprob (Brier) @ q4_under51593+0.0019[-0.001, 0.005]underpowered
MLB (in-game)CRPS @ home_margin|end_inning_320+0.4499[-0.021, 1.069]underpowered
MLB (in-game)CRPS @ home_margin|end_inning_526+0.8555[0.326, 1.489]underpowered
MLB (in-game)CRPS @ home_margin|end_inning_622+0.9583[0.241, 1.805]underpowered
MLB (in-game)CRPS @ home_margin|end_inning_721+1.0660[0.266, 1.960]underpowered
MLB (in-game)CRPS @ home_margin|end_inning_821+1.7940[0.872, 3.059]underpowered
MLB (in-game)CRPS @ total_runs|end_inning_349+0.0891[-0.303, 0.506]underpowered
MLB (in-game)CRPS @ total_runs|end_inning_554+0.2604[-0.053, 0.575]underpowered
MLB (in-game)CRPS @ total_runs|end_inning_654+0.6392[0.317, 0.963]model sharper provisional
MLB (in-game)CRPS @ total_runs|end_inning_755+0.7327[0.413, 1.081]model sharper provisional
MLB (in-game)CRPS @ total_runs|end_inning_849+1.4201[0.966, 1.898]model sharper provisional
MLB (pregame)CRPS (total_runs) @ pregame300-0.0523[-0.166, 0.048]underpowered
Soccerwinprob (Brier) @ home_win_prob|minute_6022-0.0931[-0.157, -0.034]underpowered
Soccerwinprob (Brier) @ home_win_prob|minute_7517-0.1075[-0.189, -0.032]underpowered

Model beats market only where verdict starts MODEL_SHARPER (still PROVISIONAL pending more data). Most rows are UNDERPOWERED (CI spans zero) or MARKET_SHARPER -- reported as-is.

forecaster/cross_sport_scoreboard.json · positive Δ = model sharper (paired) · a sharper verdict needs the CI clear of 0 AND enough n · edge_claimed: false

How a game moves

The aggregate transition explorer.

No committed per-game trajectory exists to replay a single match honestly, so this is the aggregate: the calibrated win-probability swing carried by each adjacent-state transition, from the state grid above. Any in-game move decomposes into the transition it crossed.

Largest calibrated win-probability swings
Soccer 15-30->30-45.6-.8->.2-.4 n>=74Soccer 15-30->30-45: -0.559 win-prob-0.559MLB mid->late.8-1->0-.2 n>=2718MLB mid->late: -0.539 win-prob-0.539MLB mid->late.8-1->.2-.4 n>=951MLB mid->late: -0.536 win-prob-0.536Soccer 30-45->45-60.2-.4->.6-.8 n>=57Soccer 30-45->45-60: +0.751 win-prob+0.751MLB mid->late0-.2->.8-1 n>=2079MLB mid->late: +0.479 win-prob+0.479MLB mid->late0-.2->.6-.8 n>=1167MLB mid->late: +0.470 win-prob+0.470
scripts/platformkit/analytics_showcase/out/why_attribution.json2026-07-23

Filtered: 13 of 20 candidate transitions are excluded because one end sits in a degenerate bucket — a realized rate of exactly 0.000 or 1.000, where every game in the bucket resolved the same way. Those buckets have no measurable swing to report; their apparent delta is thin one-sided support, not a move. Rows are also de-duplicated by destination state so one to-state cannot render as twin bars. What remains is what the data actually supports, which is mostly MLB and soccer.

Future exhibit · not yet published

Replay: watch the model think. Step through one real game’s states with the win probability and its attribution at every tick, receipts attached. This needs committed per-game trajectory data, which we haven’t published yet — so it is marked pending rather than mocked up. We never fabricate a trajectory.

Ask Scout about the forecaster
Does conditioning on the in-game state actually sharpen the forecast?Does the model actually beat the betting market?Which sport has the widest model-vs-market calibration gap?Ask anything →