Explainer
The Graveyard: Why We Publish What Failed
A search that only ever confirms is not a search. Nulls outnumber confirms here by about two to one, and every one is kept.
The most persuasive thing in this system is not a winning number. It is the pile of losing ones, kept on purpose. Across the four sport validation ledgers, nulls (351) outnumber confirms (168) by about 2.1 times. Broken out, NBA is 45 confirmed against 31 null, soccer 29 against 24, tennis 24 against 20 — and MLB actually inverts, 32 confirmed against 37 null. A separate interaction factory, which composes confirmed single findings into two-way candidates, adjudicated 1,003 of them and landed 38 confirmed, 239 null, and 709 not-testable. A system that only ever confirmed would not be credible; this one publishes the shape of a real search.
The reject ledger is the graveyard at scale: 804 recorded verdict rows, 627 of them explicit REJECTs, which resolve to 68 distinct sport-signal pairs still sitting on a reject verdict at their latest test. The gap between 804 and 68 is disclosed openly — the same dead ideas get re-tested by repeated as-of reclaim sweeps (510 of the reject rows), not 804 separate failures. A REJECT is market-efficiency evidence, not an embarrassment, and no dollar claim rides on any of it.
Four headline numbers deserve their own tombstones, because the same person who published them also built the instruments that took them apart.
The pregame return headline was a market-follow artifact. An early grader chose its bet direction from the market's own devigged favorite and never actually read the model — the evaluation file had no prediction column — while pricing every bet at a flat line real books do not offer, with filters tuned in-sample on the same data. It was betting the market's favorite and calling the result a model edge. Read against real closing lines, the model's own unfiltered number is about -2.00 percent, break-even-minus-vig.
The end-of-third-quarter win-probability headline had a fourth-quarter feature leak: two features were computed from Q4 data, so a model that predicts Q4 was peeking at it. After removing them, the honest leak-free walk-forward Brier is about 0.141, and it is framed as a leak caught in the pipeline, not a competitive number.
A steals-and-blocks prop model reported a training R-squared of about 0.79 that collapsed to about 0.06 on a leak-free holdout — textbook leakage. The corrective regularization is now hard-coded so it takes precedence over the stale tuned parameters and the mistake cannot silently return.
And the strongest surviving candidate, an apparent assists edge, was retracted on 2026-07-21 as regime-dependent: it broke in the playoffs, confirmed by an in-series replay. Under the no-edge rail, no dollar or ROI edge is claimed anywhere; the historical measurement stays in the gate artifacts only as a record of the method.
The through-line is the honest, correct result for an efficient market: against real closes the model is break-even-minus-vig, every candidate edge was rejected or retracted by its own gates, and the negative results were written down instead of quietly deleted. That is the transferable skill on display — not building an ambitious system, but building the harnesses that disprove your own best claims, and keeping the tombstones where anyone can read them.
Sources
docs/evidence/retraction-story.mdwebapp/public/data/ask/system-honesty.jsonwebapp/public/data/ask/calibration-market.jsonNext explainer
How AI Agents Built This, and What a Human Decided →