Measurement workspace
Current public page built from staged chart artifacts and a static site manifest.
CourtVision evidence
Browse the public chart artifacts first, then follow their source modules and documentation in the public repository. The market-efficiency finding and documented limitations remain part of the record.
Published snapshot
Counts are read from the committed site manifest, not a live service.
Architecture narrative
Capability map
Current public page built from staged chart artifacts and a static site manifest.
Repository architecture and limitations are documented in the linked public records.
Not represented here as an operating public service or live-status claim.
Future analytics require their own published evidence before they appear in this gallery.
Evidence gallery
Static artifacts from the public manifest. This page does not report a live backend state.
56 published charts shown
Published chart; see the linked source module for its recorded method.
Per game, order score ticks by (game-clock, ts); a margin threshold's point of no return is the first tick from which |home_score-away_score| stays >= threshold to the end. Reads only the score path, not model/market prob.
mlb: BSS(model vs market) in [-0.113, -0.092] -- model does not beat the closing reference forecast; BSS(model vs climatology) in [+0.081, +0.621] (BSS>0 means it beats that baseline). soccer_intl: BSS(model vs market) in [-1.171, -0.469...
Brier scored only for market types whose corpora carry resolved outcomes: mlb_moneyline: model Brier 0.1668 vs market 0.1510 (n=27351); soccer_match: model Brier 0.3325 vs market 0.1952 (n=4265). Unscorable market types (no resolved outc...
Published chart; see the linked source module for its recorded method.
mlb/model: 3/10 eligible bins have a calibration-gap 95% CI that excludes 0; 7 within noise. mlb/market: 0/10 eligible bins have a calibration-gap 95% CI that excludes 0; 10 within noise. soccer_intl/model: 8/9 eligible bins have a calib...
Pace is a real but MODEST variance lever: across the observed NBA pace range (~92-104) the same matchup's upset prob shifts by only a few points -- fewer possessions favor the underdog (sqrt(N) scaling), but the effect is small because r...
Descriptive only
Published chart; see the linked source module for its recorded method.
Model beats market only where verdict starts MODEL_SHARPER (still PROVISIONAL pending more data). Most rows are UNDERPOWERED (CI spans zero) or MARKET_SHARPER -- reported as-is.
team win rate with player ACTIVE vs MISSED, inside the player's own tenure window; Wald 95% CI on the win-rate difference. RAPM-free proxy.
halftime margin & second-half margin per team-game from running pbp scores; split each team's games by led/trailed at half
volume/coverage stats only -- no accuracy/quality claims (MOT metrics unbenchmarked, see JOB_EVIDENCE_PACKET s4)
Published chart; see the linked source module for its recorded method.
Per claim family (sport x hypothesis) from the 4 validation ledgers: verdict history ordered by run_ts, current status = last entry. Families with >1 distinct status flagged flipped.
nulls (351) outnumber confirms (168) 2.1x -- we publish our nulls
Published chart; see the linked source module for its recorded method.
This table's only cross-sport-comparable number is the Murphy reliability component, for 2 of 4 sports where a market-side probability decomposition exists (mlb moneyline, soccer_intl moneyline). CRPS (mlb totals/margin), Brier-without-d...
Published chart; see the linked source module for its recorded method.
mlb: at largest disagreement (>=.10, n=8530), market usually right (model_closer_rate=0.199, model_brier=0.1778 vs market_brier=0.1298). soccer_intl: at largest disagreement (>=.10, n=2291), market usually right (model_closer_rate=0.105,...
Published chart; see the linked source module for its recorded method.
Every named mechanism tested, with its recorded verdict and evidence pointer. NULL, REJECT, and NOT_TESTABLE are market-efficiency evidence. This is a descriptive measurement artifact.
Survival rate = share of TESTABLE hypotheses whose latest recorded verdict is CONFIRMED_LOCAL / CONFIRMED_LOCAL_incl_2026_OOS / REPLICATED. A NULL or REJECT is honest market-efficiency evidence.
pregame per-series consecutive-snapshot |devigged_prob| movement, bucketed by minutes-to-commence, from own scraped line_history
wnba: Brier sharpens toward the close -- T-6h=0.2454 -> close=0.2280 (delta=+0.0174, n_games=30). soccer_intl: Brier does NOT sharpen toward the close -- T-24h=0.1355 -> close=0.1374 (delta=-0.0019, n_games=13) (UNDERPOWERED).
Published chart; see the linked source module for its recorded method.
soccer_intl completes half its pre-game line motion by 2.7h before tip; nba not until 4.8h. tennis, wnba move most of their line >6h before tip.
mlb stays contested later (LCF 0.737) than soccer_intl (LCF 0.599). (nba comeback-rate cross-check 0.0146)
Most one-star-fragile: DEN (Nikola Jokić, delta_winprob 0.5822). Two estimators name the same #1 player for 1/30 teams (directional cross-check, different seasons).
mlb: MFP rises from 0.05 to 0.41 over the game (mean 0.15); soccer_intl: MFP rises from 0.68 to 1.49 over the game (mean 0.92)
soccer_intl: overshoot 0.286 and model is closer only 22% of the time at maximum disagreement, yielding OHG 0.081 as a measured failure to harvest; mlb: overshoot 0.072 and model is closer 38% of the time at maximum disagreement.
Most-taxed schedule: DEN 2025-26 (-0.38 pts/100 ORtg season-averaged); least: MEM 2025-26 (-0.27). League schedule-inequality range 0.11 pts/100 across 90 team-seasons.
Published chart; see the linked source module for its recorded method.
Support = raw input-column coverage only, never a claim that a branded metric is reproduced.
answer engine: 87/87 regression-bank checks pass (fail-closed statuses graded as PASS); honest coverage 36.6% of answerable questions -- refusals (no_data/not_supported/ambiguous/refused) are the fail-closed design, not a defect
A REJECT or DEFER is honest market-efficiency evidence.
Published chart; see the linked source module for its recorded method.
mlb: median within-game lag-1 residual autocorr = model 0.971 (n=173 games), market 0.965 (n=173); share>=0.9 model 0.90 vs market 0.86. soccer_intl: median within-game lag-1 residual autocorr = model 0.966 (n=26 games), market 0.952 (n=...
soccer_intl corpus: n=4265 rows -- far smaller than mlb/nba, wider CIs, single-fold reads not durable. Murphy: model Brier 0.3325 vs market 0.1952 (gap=+0.1374); market leads. Minute-bucket ECE (weighted): 0.3748.
n=153 shared national teams. Spearman rho=0.6419, Pearson r=0.6369 between trailing-10 form win-rate and net-xG strength -- a parallel-forms rank-stability read, NOT a temporal split-half (that series is absent from the store) and NOT a ...
Published chart; see the linked source module for its recorded method.
tennis has a served live model but no ingame_grade_joined corpus locally -- this showcase surfaces the two REAL preregistered gate receipts that ARE present (pregame cross-corpus + ingame surface-context detail layer), not a recomputed i...
market-microstructure MEASUREMENT (latency/cadence) -- NOT a trading signal
Largest calibrated win-prob drop: mlb mid(inn4-6):.8-1 -> late(inn7+):0-.2 = -0.8985 (min support n=1462). Any in-game move decomposes into the adjacent-time bucket transition it crossed.
Same methodology, per-sport n honest. nba: market ECE=0.006 (n_games=1593), fav_gap=+0.005/dog_gap=-0.001 (favorite-longshot-consistent), comeback~0.015. Not buildable: mlb, soccer, tennis.