Demo -- read-only snapshot of the live system as of 2026-07-16T01:45:31Z; no live data, paper units only.
Skip to main content

evidence / Engineering depth

Connect Your Own Claude to My Forecaster -- Live MCP Demo

Every number below arrives inside a fail-closed MCP envelope with a verdict, an n, a p-value, a source_artifact, an as_of date, and edge_claimed: false. The AI cannot improvise -- it can only relay what the engine returns or say no_data. The single truth-source for any figure is docs/JOB_EVIDENCE_PACKET.md. The three exchanges below were captured live on 2026-07-22 and are quoted verbatim.

strongest single receipt

Any Claude -- Claude Code, Claude Desktop, or an SDK agent -- can connect to this system's MCP server and get receipt-backed answers

the claim

Any Claude -- Claude Code, Claude Desktop, or an SDK agent -- can connect to this system's MCP server and get receipt-backed answers. The server does not hand the model a paragraph to paraphrase. It hands back a structured envelope: a status the model must honor verbatim (ok / no_data / not_supported / refused / ambiguous), a verdict, sample sizes, p-values, the exact file the number came from, and the snapshot date. There is no room for the model to round up, borrow a stale number, or invent a dollar edge. When the data is absent, the honest answer is no_data, and the model is instructed to say so rather than fill the gap. Below are three real exchanges. Each one shows a different reason this matters.


receipt: see evidence page source

committed artifact
docs/evidence/mcp-live-demo.md

why this matters

Fail-closed answer engines are exactly what AI-engineering teams are trying to build right now: an LLM that can only relay validated facts, cannot hallucinate a number, and admits when it doesn't know. Retrieval-augmented chat usually means "the model paraphrases some documents and hopes." This is the harder version -- every tool returns a typed envelope, the model is contractually bound to honor the status, the numbers carry their own provenance and sample size, and the honesty rails (edge_claimed: false, the caveat ladder, the market-sharper disclosure) are enforced by the server, not the prompt. The demo isn't that the model gives good answers. It's that the model cannot give a dishonest one -- and when the market beats the model, the model is the first to tell you. That fail-closed contract, not any single forecast, is the transferable engineering.


no edge claimed

This site reports calibration and sharpness only, never a dollar edge, ROI, or bankroll result. An honest null is a success.