Research papers
Studies of forecast reliability and game context
Each paper states a question, names the committed artifact behind every number, and ends on what the measurement does not establish. Calibration and description only.
Start here
A short route through the research record
Read the paper format first, then the calibration record, then the testing record.
- How to read a CourtVision paperThe contract behind every printed number, and what a paper is not allowed to claim
- Where forecasts match outcomesReliability of CourtVision win forecasts against the closing reference, by sport and by game state
- How proposed signals survive testingThe claim ledger, verdict changes and mechanism survival, as a record rather than a story
29 of 29 papers
Age and production, lightly
Why the NBA aging-curve check returned not_buildable, and what a cross-sectional profile would not have shown
An aging curve plots average production against player age. The aging_curve_lite module was asked to build one from the CourtVision NBA corpus and returned a status of not_buildable. Its gate needs two things at once:...
Catchers, umpires and hard contact
Descriptive Statcast leaderboards, their sample-size floors, and why a leaderboard is not a forecast
This note reads three published MLB Statcast-derived leaderboards side by side: catcher and umpire out-of-zone strike rate -- the artifact's 'called/swung-strike rate', which its label says is NOT a called-strike or...
Comebacks by deficit and time
How often an NBA side that is behind at a given moment goes on to win, by deficit size and time remaining, with every cell's support
This paper measures how often a side that is behind at a given moment in an NBA game goes on to win, broken out by how large the deficit is and by how much game time remained when that deficit was observed. The...
Does calibration hold up over time and across market types?
Monthly and per-market-type reliability of CourtVision forecasts against the closing reference
We checked whether CourtVision's in-game calibration holds steady across calendar months and across market types, using three showcase artifacts rebuilt on 2026-09-16 from the segment-clean in-game corpus:...
Five-man lineups, on/off and the active/missed proxy: what NBA lineup evidence can and cannot say
Reading published lineup-synergy residuals, on/off net rating and a Wald active/missed win-rate proxy as description, not forecast
This paper reads four published NBA lineup measurements side by side: a five-man lineup-synergy residual, an individual on/off net-rating leaderboard, a team win-rate active/missed proxy, and an on/off rim-deterrence...
Form curves and consistency: what rolling NBA production shows, and what it does not forecast
Reading 10-game rolling composites, shrunk consistency profiles and a pairwise scoring grid as description, not prediction
This paper reads three published measurements of NBA box-score production, pooled over the 2023-24 through 2025-26 seasons, and asks what each one does and does not establish. The form metric is a 10-game...
Halftime margins and the second half
What running-score halftime margins say about second-half margins in the NBA, and why the per-team version is masked
This paper measures how a team's second-half scoring margin relates to its halftime scoring margin in NBA games, using running play-by-play scores recorded in ctx_team_states.json (source_artifact: data/nba_ai.db ::...
Held-Out Brier, Diebold-Mariano, and the Myths That Did Not Survive
Tennis forecasts tested against their own reference model, not the market close, plus three preregistered beliefs that came back null
CourtVision publishes two tennis calibration receipts and one preregistered myth scoreboard, and this paper reads all three straight -- including one correction the title itself needs. Neither tennis_showcase.json block...
Home and away, taken apart
What the home_away_anatomy artifact decomposes for the NBA, why it never publishes a home win rate, and why it cannot be read against another sport's venue artifact on one axis
The home_away_anatomy artifact is a player-game decomposition of NBA box-score production by venue, not a decomposition of game outcomes and not a cross-sport comparison. It pools 74,450 valid player-game rows across...
How accurate is the reference forecast?
Devigged accuracy by source, the favorite-longshot shape, and how much the close sharpens over pregame snapshots
This paper checks the reference forecast against itself: how much closing devigged prices from different sources actually differ, whether the classic favorite-longshot miscalibration pattern shows up in those prices,...
How many games actually sit behind a bin
Within-game autocorrelation, effective sample size, and the game-cluster bootstrap for MLB and international-soccer calibration
Two published calibration exhibits report reliability bins built from thousands of in-game probability rows, but a row is not a game. Within a match the probability paths move in small steps from one tick to the next,...
How Much History Carries Over
Tennis surface transfer and international-soccer form stability, measured with their support
How much of a tennis player's surface-specific record, or a national soccer team's recent run, carries into a different slice of its own history? We answer this with two descriptive-only snapshots....
How proposed signals survive testing
The claim ledger, verdict changes and mechanism survival, as a record rather than a story
CourtVision keeps a validation ledger of candidate signals across basketball, baseball, soccer and tennis, and republishes that ledger rather than only the signals that held up. This note reads one dated snapshot of it...
How sample size changes a rate leaderboard
Empirical-Bayes shrinkage of three MLB rate leaderboards in a fixed 2022-2023 window
This note examines three MLB rate leaderboards from a fixed 2022-2023 Statcast corpus slice (observation_window.corpus_id statcast_fuller_v1, observation_window.as_of 2026-07-05, module mlb_shrinkage, as_of 2026-07-24)...
How the closing reference moves
Foresight by checkpoint, movement half-life, pregame absorption, and overreaction by bucket -- measured, not explained
This paper measures four properties of the devigged closing price, the reference forecast used across the site, without attributing any of the movement to a cause. First, a market foresight premium (MFP) series shows...
How to read a CourtVision paper
The contract behind every printed number, and what a paper is not allowed to claim
This paper describes the research-paper section itself and demonstrates the contract every other paper in it follows. A paper is a short method note: it states one question, names the committed artifact and field path...
How unequal are the leagues?
A parity index across sports, with its window and its floors
This paper reads the CourtVision league_parity_index module, which measures competitive balance within a league across seasons using teams' win-share concentration (a Gini coefficient and a Herfindahl index) and...
Pace, star removal and momentum in the NBA
What a counterfactual simulator says a variance lever is worth, what removing a lead player costs a modeled matchup, and which momentum-shaped claims survived testing
This note pulls together three NBA measures: a schedule-legal variance lever, a modeled lineup subtraction, and preregistered momentum tests. The pace-variance simulator fixes a matchup's strength and per-possession...
Pitch velocity shape: what the distribution says, and what it does not
Per-type velocity percentiles and mph bands across 19 published 2025 Statcast pitch-type codes
This note describes the shape of pitch velocity by pitch-type code in a local 2025 Statcast pull (source data/cache/statcast/statcast_fuller__2025.parquet), not just its average. statcast_showcase.json publishes 19...
Rest is a relative quantity
The NBA schedule differential tracks outcomes, the shared-congestion groups look alike, and the one season with a recorded reference forecast cannot tell us whether either is carried
Two claims about NBA rest are usually run together: that a tired team plays worse, and that a team tired relative to its opponent wins less often. The schedule separates them. Over 4,793 regular-season games from...
Rest, load and outcomes
The schedule-fatigue tax and the load-bearing index in the NBA, with their observation windows kept separate
This note describes two separate NBA descriptive measures and keeps their observation windows apart rather than blending them into one story. The Schedule Fatigue Tax (SFT) multiplies each of 90 team-seasons' own...
The graveyard: what the null results teach
A taxonomy of why proposed signals do not clear the gate, their recorded frequencies, and the case for publishing nulls
CourtVision keeps two published ledgers of ideas that did not ship, read here as a record of a search process, not as findings about any one idea. honesty_exhibit.json, generated 2026-07-22, buckets 1290 tested rows...
The rotation day that was already known
MLB starting-pitcher rest moves the win frequency by less than the noise, and the recorded closing reference forecast already carries whatever is left
Baseball has argued about the fourth and fifth day of a starting pitcher's rotation for as long as there have been five-man rotations. This note measures it twice on the same corpus: once on its own, and once against...
What carries across sports
Which reliability comparisons are supported across NBA, MLB, soccer and tennis, and what a shared metric scale does not establish
CourtVision runs the same reliability decomposition and the same skill-score formula in more than one sport, but an identical formula run in two places does not by itself make the two results comparable....
What the Brier score is made of
Reliability, resolution and uncertainty in CourtVision's MLB and international-soccer forecasts, checked against the closing market
CourtVision splits each in-game win-probability Brier score into three Murphy components -- reliability, resolution and uncertainty -- and stores both a directly computed Brier and a reconstructed Brier evaluated on the...
What the count does to the next pitch
Conditional next-pitch probabilities and count-class pitch mixes in MLB, 2025 Statcast
This note compares two descriptive views of how the ball-strike count shapes a Major League pitcher's next pitch, both built from a local 2025 Statcast pull (source data/cache/statcast/statcast_fuller__2025.parquet, as...
When a lead lasts
Incidence and timing of margins that persist to the last recorded score tick, MLB and international soccer
This paper asks a narrow, descriptive question about game flow: once a scoring margin opens to a given size, how often does it hold all the way to the last score tick this system recorded, and when in the game does that...
Where forecasts match outcomes
Reliability of CourtVision win forecasts against the closing reference, by sport and by game state
We measured how closely CourtVision's model win probabilities and the devigged market close match observed outcomes, in aggregate and by game state, on two segment-clean in-game corpora: MLB (27,351 rows, 178 games) and...
Where the forecast errors live
Residual anatomy of MLB and international-soccer win forecasts, by game-state time bucket and forecast-probability band
We decompose forecast error for two segment-clean in-game corpora -- MLB (27,351 rows, 178 files) and international soccer (4,265 rows, 27 files) -- into cells crossing a game-state time bucket with a five-band...