Skip to content

Findings

The honesty exhibits

These pages make the honest rail concrete. No dollar edge is claimed anywhere; each exhibit either takes a number apart, deflates our own sample, or publishes a result that failed — on purpose.

Retractions

The six numbers we took back, each with what it was, why it was wrong, and where it now lives only inside its retraction.

Effective sample size

We deflate our own row counts: 78,986 within-game MLB rows carry the independent information of at most ~227 games, so confidence intervals must widen accordingly.

Verdict flips

The claim families that changed their mind as more data arrived -- each mind-change in sequence, plus how long the retracted claims lived. A flip is the process working.

MLB leaderboards, with the nulls

Descriptive Statcast-derived leaderboards from a fixed 2022-2023 window, published beside the gate nulls that failed -- park factor too unstable to rank, umpire tendency does not move totals.

Tennis: momentum's grain limit

Momentum is real point-to-point (n=456,383, confirmed) but dies at the game grain; two tiebreak myths come back null; altitude and travel effects are confirmed, altitude replicated across two disjoint year slices.

NBA momentum, tested

The honest split: structural fatigue, rest, and clutch effects are confirmed and some replicated, while the individual hot/cold carryover shapes -- even a player's own B2B dip at n=65,103 -- come back null.

The fourth-quarter shift

The module that honestly refused a clutch number, turned into a descriptive one: per-player Q4-vs-earlier per-36 shifts over 1,231 games -- published with the blowout/garbage-time confound in plain sight, never as a clutch or predictive claim.

When the leaderboard regresses

Empirical-Bayes shrinkage on the MLB rate leaderboards: small-sample 'leaders' get pulled toward the group mean by how few trials back them, while a huge-sample rate barely moves. The raw-vs-shrunk gap is the honesty.

Are our probabilities honest?

Calibration reliability diagrams and the Murphy Brier decomposition: our model's Brier trails the market's (0.2377 vs 0.2067 in MLB), and the gap is resolution -- information we don't have -- not miscalibration. We match or trail the close; we never claim to beat it.

The life of a forecast

How a market absorbs information, twice: pre-game line half-life per sport, then in-game Brier checkpoints where the market stays ahead of our model at every point measured. Descriptive only, same honest gap as reliability.

Home advantage, decomposed

49,425 international matches split on the real neutral-site flag: the true-home/neutral goal-diff gap grew from 0.22 to 0.50 not because true-home advantage rose (flat) but because the neutral-venue edge collapsed over time.

We graded the bookmakers

Proportional-devig Brier on shared-game subsets: Pinnacle is barely the sharpest book on tennis match-winner (0.1975 vs Bet365's 0.1980), while on soccer over/under 2.5 goals the books are a statistical dead heat.

Who bends the shot chart

A transparent on/off rim-deterrence leaderboard from real NBA zone splits: Wembanyama (-0.0618) and Gobert (-0.0632) top their seasons, matching the eye test -- published beside the roster confound it can't remove.

How competitive is each season?

A 3-season parity ledger: win-share Gini rose from 0.177 (2023-24) to 0.2004 (2025-26), published beside the partial-season and playoff-pooling caveats that keep it a description, not a trend.

Greater than the sum of their parts

A five-man lineup-synergy ledger: the Grizzlies' Jackson-Morant-Bane-Edey-Wells five topped 2024-25 at +24.32 net-per-48 above expected -- published beside the single-season and small-minutes confounds.

Is the market calibrated?

A favorite-longshot bias audit that grades the MARKET, not our model: tennis shows a mild, monotone bias favoring favorites (gap grows from -0.0026 to +0.0176 across five buckets), while MLB moneyline is essentially efficient.