Skip to content
The self-improving system

An AI that grades itself.

It proposes signals, tests them under a leak-free gate, and keeps every null, not-testable, and retraction in the same ledger it confirms from. A system that only ever confirmed would not be credible — so it counts what it could not prove, and changes its own mind on new evidence.

confirmednull (a finding)not testabledescriptivepending

The scoreboard grades every claim

259
claim families tracked
121
verified (confirmed)
134
null, not-testable, or retracted
5
changed their own verdict

The scoreboard above grades forward-tested claim families; the mechanism ledger is a separate, wider census of every named mechanism. Across 287 named mechanisms: 130 confirmed, 126 null, 31 not-testable.

nulls (351) outnumber confirms (168) 2.1x -- we publish our nulls

The system changes its own mind

These families carried more than one verdict across reruns. The ledger keeps the whole sequence, not just the last word — the trail of a system re-testing itself as corpora grow.

NBAthree in four fatigueNULLCONFIRMED
MLBreliever 3in3d fatigueNULLNULLCONFIRMED
MLBreliever 3in3d fatigue combinedPROVISIONALNULL
MLBstaff dayafter fatigue chainNULLCONFIRMEDNULL
Tennisbreak point conversion by set numberNOT TESTABLECONFIRMED

The graveyard is a feature

Every reject and defer is recorded, categorized, and counted. A discarded signal is honest market-efficiency evidence, not a failure to bury.

628
reject rows recorded
68
distinct signals on a reject verdict now
804
total verdict rows across all reruns

Top reject reasons: 510 asof reclaim sweep; 20 prescreen delta; 18 wf ablation · cumulative reaches 628 by 2026-07-18.

Ask Scout about the loop
How does this system grade itself?How many candidate market signals were rejected?Do any of your findings flip verdict when you re-test them?Ask anything →