Skip to content

Findings / Retraction

The Retraction Story

The most persuasive thing on this site is not a winning number — it is the pile of losing ones, kept on purpose. These six headline figures were each published once, then taken apart by the same instruments that built the system, and every replacement below is calibration-only: no dollar, ROI, or edge figure is claimed anywhere on this site.

The single truth-source for every figure below is docs/JOB_EVIDENCE_PACKET.md (packet as_of 2026-07-23). These six numbers appear here, and only here, inside this retraction framing — see What verified means for the discipline that produced this table.

Retracted

+18.38% pregame ROI on 1,535 walk-forward bets vs real closing lines

What was wrong

Market-follow artifact, confirmed at the source-code level. The grader picked bet direction from the market's own devigged lean and never read the model (the eval CSV had no prediction column), priced at a flat -110 that real books do not offer, and tuned its filters in-sample on the same file.

Proof artifact

JOB_EVIDENCE_PACKET s4 (model's own unfiltered number: -2.00%)

Honest replacement

Roughly break-even-minus-vig vs real closing lines. Every candidate edge, including assists, was ultimately rejected or retracted by the same gates.

Retracted

0.119 end-of-Q3 in-play Brier, "inside Pinnacle's range"

What was wrong

Leak-inflated and mis-sourced. Two features were computed from fourth-quarter data, so the model predicting Q4 was peeking at Q4; the cited file actually reported 0.1354, a different number.

Proof artifact

JOB_EVIDENCE_PACKET s3 (leak-free re-run; controlled A/B ~4% relative inflation)

Honest replacement

Leak-free walk-forward end-of-Q3 Brier ~0.141, after removing the Q4 feature leak found in the pipeline. Framed as a leak caught, not a competitive number.

Retracted

+54.57% ROI / 78.11% hit on 55,073 in-play bets

What was wrong

Graded against an L5 line proxy, not real closing lines. A model-quality ceiling on a soft proxy, never a tradeable result.

Proof artifact

JOB_EVIDENCE_PACKET s4

Honest replacement

On a soft L5 proxy the in-play backtest reaches that ceiling. Treated strictly as a model-quality ceiling, never as realized edge.

Retracted

Aggregate CLV +8.94pp

What was wrong

Circular -- computed on the same model-unused, devig-direction corpus. No real Pinnacle-close CLV exists yet; a full-season backtest shows CLV about zero vs real closes.

Proof artifact

JOB_EVIDENCE_PACKET s4 (full-season backtest: CLV ~= 0 vs real closes)

Honest replacement

Real closing-line CLV cannot be measured yet. The methodology that will measure it exists; no CLV figure is quoted until it can be.

Retracted

Steals/blocks prop grid search: training R^2 ~0.79

What was wrong

Textbook leakage: the ~0.79 training R^2 collapsed to ~0.06 on a leak-free holdout.

Proof artifact

src/prediction/prop_cv_split.py (documents the gap; hard-codes corrective regularization)

Honest replacement

Caught and hard-corrected a leakage-driven overfit. The corrective regularization takes precedence over the stale tuned parameters, so the mistake cannot silently reappear.

Retracted

The assists ROI edge (the strongest surviving candidate)

What was wrong

Regime-dependent -- it broke in the playoffs -- and retracted 2026-07-21. Under the no-edge rail, no dollar/ROI edge is claimed anywhere.

Proof artifact

JOB_EVIDENCE_PACKET s3 (historical record only, in the gate artifacts)

Honest replacement

No dollar or ROI edge is claimed anywhere. The historical measurement remains only as a record of the stress-testing methodology.

The through-line: against real closing lines the market is efficient, the model is break-even-minus-vig, and every candidate edge, including the strongest one, was rejected or retracted by its own gates. That is the honest, correct result for an efficient market — and the harnesses that prove it are the same ones that took these six numbers apart.

Ask Scout about this
What was the biggest result you had to retract?Do you keep retracted findings in the ledger or delete them?How many of your claims are verified versus null?Ask anything →