Skip to content

Findings

Findings / Data integrity

MLB in-game calibration now runs on a segment-clean corpus.

This check, measured 2026-09-16, found that a portion of the MLB in-game corpus joins ticks from consecutive-day games and applies one outcome label across them. The corpus has since been re-segmented so that one stored file holds one real game, and the affected artifacts were rebuilt from it. The published MLB and international soccer numbers are now revision 2; the revision 1 values measured here are withdrawn and kept in the regeneration receipt.

Receipt v1 / published incident receipt / regeneration receipt / timing regeneration receipt / scripts/platformkit/check_ingame_join_integrity.py

What was measured

SportFilesTicksMixed filesMismatched ticksNo game stateTruncated gamesLabel disagreements
MLB22778,98612627,07626,340610
International soccer519,0031Not reported4,64922 (Draws truncated at minute 60 and 42.)

Scroll horizontally for all columns.

Game-level labels agree with the last stated score in all 166 games with a leader (frequency 1.0000).

Tick-level agreement

The comparison is limited to late innings, defined here as inning 7+ and to ticks where a leading side can be assessed.

PopulationLeading-side label agreementSupport
Late-inning (inning 7+) leading-side labels across all MLB ticks0.7284n = 12,588
Late-inning (inning 7+) leading-side labels in the final MLB segment0.9200n = 6,623

Scroll horizontally for all columns.

What the regeneration changed

Every stored tick file was re-segmented so that one file holds one real game, the segmented corpus was re-checked with the same fail-closed integrity checker, and the 13 exposed artifacts were rebuilt from the segmented corpus. Every figure below is the published headline of one artifact before and after that rebuild, measured 2026-09-16 on data/cache/ingame_grade_joined/{mlb_segmented,soccer_intl_segmented}.

ArtifactPopulationn beforen afterHeadline before (revision 1, withdrawn)Headline after (revision 2)
state_conditioned_calibrationMLB ticks bucketed78,98627,351MLB n-weighted calibration error: model 0.079, reference quote 0.0591MLB n-weighted calibration error: model 0.0494, reference quote 0.0397
calibration_stabilityMLB ticks resampled78,98627,351MLB reliability bins whose gap interval excludes zero: model 4 of 10, reference quote 0 of 10MLB reliability bins whose gap interval excludes zero: model 3 of 10, reference quote 0 of 10
murphy_decompositionMLB ticks decomposed78,98627,351MLB model reliability 0.012701, resolution 0.023628MLB model reliability 0.005619, resolution 0.087187
brier_skill_scoresMLB ticks scored78,98627,351MLB Brier skill score against the reference quote, all grains: -0.1502MLB Brier skill score against the reference quote, all grains: -0.1043
residual_anatomyMLB ticks ranked78,98627,351largest segment absolute-residual mass 6063.25 over 12344 rows (MLB early(inn1-3), .4-.6)largest segment absolute-residual mass 2812.37 over 5754 rows (MLB early(inn1-3), .4-.6)
calibration_by_market_typeMLB moneyline ticks scored78,98627,351MLB moneyline Brier: model 0.237684, reference quote 0.206653MLB moneyline Brier: model 0.166759, reference quote 0.151008
residual_autocorrelationMLB ticks in series78,98627,351MLB median within-game lag-1 residual autocorrelation: model 0.9809 over 222 gamesMLB median within-game lag-1 residual autocorrelation: model 0.9713 over 173 games
calibration_over_timeMLB ticks split by month78,98627,351MLB 2026-07 Brier: model 0.2357, reference quote 0.1938MLB 2026-07 Brier: model 0.16, reference quote 0.1428
calibration_atlas (not published as a module)ticks in the largest atlas entry18,6119,04226 atlas entries; largest entry backed by 18611 ticks26 atlas entries; largest entry backed by 9042 ticks
market_disagreement_profileMLB ticks in the widest disagreement bucket33,4028,530at disagreement >=.10 the model is closer 0.3773 of the time (Brier 0.282705 vs 0.210258)at disagreement >=.10 the model is closer 0.1994 of the time (Brier 0.177763 vs 0.129824)
info_arrival_curveMLB ticks at the 9th-inning checkpoint2,2821,110reference-quote-minus-model Brier widens to -0.0902 by the 9th inningreference-quote-minus-model Brier narrows to -0.001 by the 9th inning
market_overreactionMLB ticks in the 0-1pt move bucket66,69719,044MLB 0-1pt moves: moved-to price minus outcome rate 0.0689MLB 0-1pt moves: moved-to price minus outcome rate 0.0302
soccer_calibration_packsoccer ticks scored9,0034,265soccer Brier: model 0.227887, reference quote 0.142726soccer Brier: model 0.332522, reference quote 0.195157

Scroll horizontally for all columns.

After segmentation both sports report 0 label disagreements and 0 multi-game files, and no tick sits outside its game's date. Two residual single-file defects remain and are disclosed rather than removed: one frozen-quote file per sport, plus 4 MLB and 1 soccer file shorter than 10 ticks.

These are population changes, not forecaster improvement. The forecaster did not change. Two different shares are easy to conflate, so both are stated here. The specifically identified label mismatches were 27,076 of the 78,986 MLB ticks, 34.3 percent. The share segmentation actually excluded is larger: MLB kept 27,351 ticks and excluded 51,635, 65.4 percent, because mixed-game files were dropped whole, stateless ticks carry no ordering information, and the remaining segmentation exclusions applied on top. Soccer kept 4,265 of 9,003 ticks and excluded 4,738, 52.6 percent; its identified-mismatch count was never reported separately. Read every move as what the measurement should have said all along.

The smaller population requires newly estimated intervals rather than rescaled ones: the cluster-bootstrap intervals published with these artifacts were re-estimated on the segmented corpus, not widened from the revision 1 figures. A lower Brier or calibration error after segmentation is not evidence of a sharper model, and none of these figures is a claim about money.

Soccer moves the other way for a mechanical reason: the draw share of scored ticks fell from 0.2615 to 0.1376, and a draw (outcome 0.5) is cheap to score against a two-way probability, so BOTH the model and the reference quote post a higher Brier on the segmented soccer sample. The model-minus-reference gap is the like-for-like line to read; it halves for MLB (0.031031 to 0.015752) and widens for soccer (0.08516 to 0.137365).

The single largest correction is the late-inning high-confidence MLB cell: the model's own forecast of 0.9211 was scored against a 0.6853 outcome rate drawn partly from other games. On the segment-clean corpus the same cell reads 0.9101 against 0.9498.

Read the full regeneration receipt

What the timing regeneration changed

The five timing artifacts left under review, plus the derived market-foresight premium, were rebuilt from the same segmented corpus the first pass used. Each producer takes its corpus root from CV_INGAME_CORPUS_SUFFIX; with the suffix set to _segmented the glob-backed producers read <sport>_segmented and the artifact-backed producers read the staged out_segmented copies, so no artifact can mix a rebuilt input with an unrebuilt one. The rebuild ran with CV_INGAME_CORPUS_SUFFIX=_segmented, measured 2026-09-17 on data/cache/ingame_grade_joined/{mlb_segmented,soccer_intl_segmented}.

ArtifactPopulationn beforen afterHeadline before (revision 1, withdrawn)Headline after (revision 2)
blowout_dynamicsMLB games with a usable score path178174MLB (178 games): a 2-run gap becomes permanent by a median inning 6.0 in 75 percent of games; a 5-run gap by a median inning 7.5 in 27 percent. Soccer (29 games): a 1-goal gap by a median minute 39.5 in 48 percent.MLB (174 games): a 2-run gap becomes permanent by a median inning 5.0 in 75 percent of games; a 5-run gap by a median inning 6.0 in 32 percent. Soccer (26 games): a 1-goal gap by a median minute 37.0 in 46 percent.
novel_live_clock_fractionMLB games behind the near-median threshold178174MLB Live-Clock Fraction 0.8333 at the 3-run threshold; soccer 0.7475 at the 1-goal threshold.MLB Live-Clock Fraction 0.7368 at the 3-run threshold; soccer 0.5994 at the 1-goal threshold.
market_convergenceMLB ticks across game-time checkpoints52,62826,677MLB disagreement between the model and the reference quote widens from 0.07 at inning 1 to 0.2791 at inning 10; reference entropy falls 0.9669 to 0.8761 bits.MLB disagreement between the model and the reference quote widens from 0.0839 at inning 1 to 0.2067 at inning 10; reference entropy falls 0.9553 to 0.8488 bits.
why_attributionadjacent-time state transitions scored120120Largest calibrated win-prob drop: soccer 15-30:.4-.6 to 30-45:.2-.4 = -0.7683 (minimum support n=120).Largest calibrated win-prob drop: MLB mid(inn4-6):.8-1 to late(inn7+):0-.2 = -0.8985 (minimum support n=1462).
comeback_atlasNBA lead-by-time state buckets848484 NBA state buckets, 7 below the n>=30-games mask.84 NBA state buckets, 7 below the n>=30-games mask -- every cell unchanged. This atlas reads the NBA reliability map built from data/cache/inplay_odds/nba_checkpoints_full.parquet, not the MLB and soccer in-game join, so the segmentation cannot move it; the re-run on the segmented roots reproduced every published cell and only stamped the corpus it actually reads.
novel_market_foresight_premiumMLB game-time checkpoints with both instruments1010MLB foresight premium rises from 0.063 at inning 1 to 0.2365 at inning 10 (mean 0.2461); soccer 0.6393 to 1.7056 (mean 0.9211).MLB foresight premium rises from 0.051 at inning 1 to 0.411 at inning 10 (mean 0.1487); soccer 0.6779 to 1.4861 (mean 0.9225).

Scroll horizontally for all columns.

Re-run on 2026-09-17 the checker reproduces the first pass exactly: both segmented roots report 0 label disagreements, 0 multi-game files and no tick outside the date of its own game, and the leading side at the last tick agrees with the label in every game. The same residual single-file defects are disclosed rather than removed: one frozen-quote file per sport, plus 4 MLB and 1 soccer file shorter than 10 ticks.

These are population changes, not forecaster improvement. No forecaster changed. Segmentation dropped 49 of 227 MLB files and 24 of 51 soccer files outright, so the timing exhibits now describe 178 MLB and 27 soccer stored games instead of 227 and 51, and 174 MLB and 26 soccer games once the 10-tick floor is applied. Read every move as what the measurement should have said all along.

Timing artifacts were hit harder than the label-consuming ones, and for a different reason. A mixed file concatenates two score paths, so a lead that never reverts inside the stored file can be two separate games' leads laid end to end. That pushes the point of no return later and inflates the share of the clock that looks contested. The MLB Live-Clock Fraction falls 0.8333 to 0.7368 and the soccer one 0.7475 to 0.5994; the median MLB inning at which a 5-run gap becomes permanent falls from 7.5 to 6.0. Those falls are the artifact being removed.

Comeback and blowout rates move in both directions because the denominator moved too. A 3-run MLB gap is now decided in 55.75 percent of games against 53.37 percent before, and a 5-run gap in 31.61 percent against 26.97 percent, while the 2-run share barely moves (74.71 against 75.28 percent). The soccer 1-goal share falls 48.28 to 46.15 percent on 26 games, which is three games of movement and not a finding.

Convergence paths are now measured on 26,677 MLB ticks instead of 52,628. The model and the reference quote still diverge over the game, but less far: the late-inning disagreement reads 0.2067 instead of 0.2791. The market foresight premium inherits both rebuilt inputs and stops mixing revisions; its MLB mean falls from 0.2461 to 0.1487 while its last checkpoint rises from 0.2365 to 0.411, which is the small late-inning sample and not a trend to read.

comeback_atlas is the honest exception. It was listed under review with the other timing artifacts, but its corpus is the NBA reliability map, not the MLB and soccer in-game join, so the segmentation could not touch it. Re-running it on the segmented roots reproduced all 84 state buckets exactly. Its status is cleared on that evidence, not on a rebuild.

The smaller population requires newly estimated intervals rather than rescaled ones: the intervals published with these artifacts were re-estimated on the segmented corpus, not widened from the revision 1 figures. For the record the two shares behind that smaller population are different numbers: 27,076 of 78,986 MLB ticks, 34.3 percent, were the identified label mismatches, while segmentation excluded 51,635 MLB ticks in all, 65.4 percent, and 4,738 of 9,003 soccer ticks, 52.6 percent. None of these figures is a claim about money.

Read the full timing regeneration receipt

Affected calibration artifacts

The receipt records these artifacts against outcome labels. Their MLB and international soccer rows were rebuilt from the segment-clean corpus and now read revision 2.

Timing artifacts regenerated on 2026-09-17: blowout_dynamics, novel_live_clock_fraction, market_convergence, why_attribution, comeback_atlas, novel_market_foresight_premium. Timing artifacts still under review: none. Mixed-game tick paths also distort timing measurements even where no outcome label is used, so these six were held back from the first pass. They were regenerated on 2026-09-17 from the same segment-clean corpus and published as revision 2; the before and after figures are in the mlb-ingame-timing-regeneration receipt. novel_market_foresight_premium no longer mixes revisions: it now reads the rebuilt info_arrival_curve and the rebuilt market_convergence. comeback_atlas was cleared on a different ground -- its corpus is the NBA reliability map, not this in-game join, so re-running it on the segmented roots reproduced every published cell. No timing artifact remains under review.

Artifact status registry

Only non-clear artifact and sport rows are listed. NBA and tennis rows remain clear.

ArtifactSportStatus
state_conditioned_calibrationMLBregenerated
state_conditioned_calibrationInternational soccerregenerated
calibration_stabilityMLBregenerated
calibration_stabilityInternational soccerregenerated
murphy_decompositionMLBregenerated
murphy_decompositionInternational soccerregenerated
brier_skill_scoresMLBregenerated
brier_skill_scoresInternational soccerregenerated
residual_anatomyMLBregenerated
residual_anatomyInternational soccerregenerated
calibration_by_market_typeMLBregenerated
calibration_by_market_typeInternational soccerregenerated
residual_autocorrelationMLBregenerated
residual_autocorrelationInternational soccerregenerated
calibration_over_timeMLBregenerated
calibration_over_timeInternational soccerregenerated
calibration_atlas (not published as a module)MLBregenerated
calibration_atlas (not published as a module)International soccerregenerated
market_disagreement_profileMLBregenerated
market_disagreement_profileInternational soccerregenerated
info_arrival_curveMLBregenerated
info_arrival_curveInternational soccerregenerated
market_overreactionMLBregenerated
market_overreactionInternational soccerregenerated
soccer_calibration_packMLBregenerated
soccer_calibration_packInternational soccerregenerated
blowout_dynamicsMLBregenerated
blowout_dynamicsInternational soccerregenerated
novel_live_clock_fractionMLBregenerated
novel_live_clock_fractionInternational soccerregenerated
market_convergenceMLBregenerated
market_convergenceInternational soccerregenerated
why_attributionMLBregenerated
why_attributionInternational soccerregenerated
comeback_atlasMLBregenerated
comeback_atlasInternational soccerregenerated
novel_market_foresight_premiumMLBregenerated
novel_market_foresight_premiumInternational soccerregenerated
novel_overreaction_harvest_gapMLBunder-review
novel_overreaction_harvest_gapInternational soccerunder-review
ess_ledgerMLBunder-review
ess_ledgerInternational soccerunder-review

Scroll horizontally for all columns.

novel_overreaction_harvest_gap is under review because it was composed from an input that has since been rebuilt. Composed on 2026-07-25 from market_overreaction and market_disagreement_profile. Both inputs were rebuilt at revision 2 on the segment-clean corpus, so every figure here is still the revision 1 composition and is under review until it is recomposed. Stale input: market_overreaction, market_disagreement_profile (not published as a module).

ess_ledger is under review because it was composed from an input that has since been rebuilt. Composed on 2026-07-25 from residual_autocorrelation. That input was rebuilt at revision 2, so this ledger still describes the joined MLB and international soccer corpus of 78,986 and 9,003 rows and is under review until it is recomposed. Stale input: residual_autocorrelation.

What clears the status

  • Done: segment identity checks keep every tick within its real game, and the checker reports no label disagreement and no multi-game file on the segmented corpus.
  • Done: the affected artifacts were regenerated from that checked corpus and published as revision 2.
  • Done: the six timing artifacts were rebuilt from the same corpus on 2026-09-17 and published as revision 2, and comeback_atlas was cleared on the separate evidence that its corpus is the NBA reliability map, which this segmentation cannot touch. No artifact remains under review.
  • Outstanding: the research papers still quote revision 1 numbers and must be re-audited against the regenerated artifacts.

Findings