Skip to content

Calibration / comparability gate

Read the published comparability decisions first.

This view shows exactly which reliability-component comparisons are supported by the committed evidence, and which measurements remain outside that axis.

Reading trail

Read first: Read each sport's calibration reliability before comparing components across sports.

Next question: How closely do published forecast bins align with observed outcomes?

Read the analysis: What carries across sports

Published measurement capability

Availability is derived from the published row fields. A null value remains not published; it is never treated as zero.

SportModel reliability scoreReference reliabilityReliability decompositionPopulation
MLBPublishedPublishedPublishedPublished
International soccerPublishedPublishedPublishedPublished
NBANot publishedNot publishedNot publishedPublished
TennisNot publishedNot publishedNot publishedPublished

Supported reliability-component comparisons

These are the only rows with model and reference reliability components on the published 0-1 probability scale.

Model reliabilityReference reliability
Comparable reliability components
Each pair retains its published population size.
0.0000.0960.192MLB model reliability 0.005619MLB reference reliability 0.003425MLBn=27351International soccer model reliability 0.191982International soccer reference reliability 0.088088International soccern=4265
public/data/showcase/kernel_transfer.jsonSnapshot generated 2026-09-17n not published
MLBmoneyline_ingame (murphy reliability decomposition)

Model 0.005619; reference 0.003425; n=27351.

View Score decomposition

Rows held outside the comparison

A common scale does not establish matched populations or transferability. These published rows do not have the compatible components required for this chart.

MLB: totals_margin (CRPS distributional, pregame+ingame)

Missing or incompatible measurement: Murphy reliability/resolution for a binary outcome probability

CRPS scores a full run-count distribution, not a binary outcome probability -- Murphy reliability/resolution is undefined for it here; never forced onto the Brier axis above.

Population n=671

View Kernel Transfer evidence

NBA: winprob_ingame (Brier, no reliability decomposition)

Missing or incompatible measurement: 10-bin Murphy reliability/resolution split

nba's benchmark stores per-checkpoint Brier MEANS only -- no 10-bin Murphy reliability/resolution split has been computed locally for nba win-prob, honest gap not an estimate.

Population n=6371

View Kernel Transfer evidence

Tennis: pregame_prior + ingame_surface_gate (no market side)

Missing or incompatible measurement: market probability

both tennis gate receipts are model-vs-model checks (prior vs base; surface-specific vs surface-blind) -- no market_prob to decompose against, reliability undefined by construction.

Population n=55075

View Kernel Transfer evidence