Findings / Shrinkage
When the leaderboard regresses to the mean
Small-sample rate leaders regress toward the group mean under empirical-Bayes shrinkage; the raw leaderboard overstates them.
How to read this: every rate below is pulled toward its group's pooled mean by an amount that depends on how few trials back it. Fit a beta-binomial prior per group (pooled mean m, concentration kappa = alpha + beta), then replace each raw rate with the posterior mean (k+alpha)/(n+alpha+beta). A player with tens of thousands of trials barely moves; a player with a few hundred can move a lot. This is a modeling choice — that the entities are exchangeable draws from one shared prior — not a newly discovered truth about any one name.
Window: 2022-2023 Statcast slice (statcast_fuller_v1) — fixed 2022-2023 slice.
catcher ooz
Catcher out-of-zone strike rate (NOT a framing/called-strike rate; descriptive)
pooled_mean 0.2813 · kappa 834.93 · alpha 234.91 · beta 600.02 · n_entities 113
Biggest regressors
The small-n rows the raw leaderboard overstated most.
| Name | n | Raw rate | Shrunk rate | Moved by |
|---|---|---|---|---|
| Drew Millas | 585 | 0.212 | 0.2528 | -0.0408 |
| Cooper Hummel | 914 | 0.2177 | 0.2481 | -0.0304 |
| Carlos Pérez | 1,277 | 0.2388 | 0.2556 | -0.0168 |
| David Fry | 882 | 0.2528 | 0.2667 | -0.0139 |
| Logan Porter | 668 | 0.262 | 0.2727 | -0.0108 |
| Brett Sullivan | 1,708 | 0.2512 | 0.2611 | -0.0099 |
| Manny Piña | 541 | 0.2976 | 0.2877 | +0.0099 |
| P.J. Higgins | 1,946 | 0.2497 | 0.2592 | -0.0095 |
| Payton Henry | 731 | 0.264 | 0.2733 | -0.0092 |
| Korey Lee | 2,159 | 0.2501 | 0.2588 | -0.0087 |
| Iván Herrera | 1,208 | 0.2624 | 0.2702 | -0.0077 |
| Carlos Pérez | 2,234 | 0.2534 | 0.261 | -0.0076 |
Scroll horizontally for all columns.
Leaderboard after shrinkage
Ranked by shrunk rate — the honest ranking.
| Name | n | Raw rate | Shrunk rate |
|---|---|---|---|
| René Pinto | 3,448 | 0.3147 | 0.3082 |
| Alejandro Kirk | 10,963 | 0.3048 | 0.3031 |
| Austin Barnes | 6,707 | 0.3049 | 0.3023 |
| Patrick Bailey | 5,657 | 0.3035 | 0.3007 |
| Travis d'Arnaud | 11,281 | 0.2994 | 0.2982 |
| Cal Raleigh | 14,504 | 0.299 | 0.2981 |
| Jacob Stallings | 12,090 | 0.2979 | 0.2969 |
| Sandy León | 2,537 | 0.3015 | 0.2965 |
| Curt Casali | 5,068 | 0.2985 | 0.2961 |
| Danny Jansen | 8,461 | 0.2975 | 0.296 |
| Jose Trevino | 9,981 | 0.2968 | 0.2956 |
| Kyle Higashioka | 9,821 | 0.2962 | 0.295 |
Scroll horizontally for all columns.
umpire ooz
Umpire out-of-zone strike rate (NOT a framing/called-strike rate; descriptive)
pooled_mean 0.2813 · kappa 2830.5 · alpha 796.18 · beta 2034.32 · n_entities 102
Biggest regressors
The small-n rows the raw leaderboard overstated most.
| Name | n | Raw rate | Shrunk rate | Moved by |
|---|---|---|---|---|
| Randy Rosenberg | 506 | 0.3142 | 0.2863 | +0.0279 |
| David Arrieta | 754 | 0.2586 | 0.2765 | -0.0179 |
| Derek Thomas | 2,608 | 0.2523 | 0.2674 | -0.0151 |
| Marty Foster | 1,492 | 0.2983 | 0.2871 | +0.0111 |
| Jacob Metz | 1,735 | 0.2986 | 0.2878 | +0.0107 |
| Brian Walsh | 1,960 | 0.298 | 0.2881 | +0.0099 |
| Greg Gibson | 1,308 | 0.2951 | 0.2857 | +0.0095 |
| John Bacon | 2,933 | 0.267 | 0.274 | -0.007 |
| Jerry Meals | 3,759 | 0.2974 | 0.2905 | +0.0069 |
| Bill Miller | 8,815 | 0.3077 | 0.3012 | +0.0064 |
| Laz Diaz | 7,197 | 0.3023 | 0.2964 | +0.0059 |
| Chris Conroy | 4,251 | 0.2955 | 0.2898 | +0.0057 |
Scroll horizontally for all columns.
Leaderboard after shrinkage
Ranked by shrunk rate — the honest ranking.
| Name | n | Raw rate | Shrunk rate |
|---|---|---|---|
| Bill Miller | 8,815 | 0.3077 | 0.3012 |
| Doug Eddings | 8,181 | 0.3023 | 0.2969 |
| Laz Diaz | 7,197 | 0.3023 | 0.2964 |
| Andy Fletcher | 8,328 | 0.3004 | 0.2956 |
| Gabe Morales | 8,409 | 0.2988 | 0.2944 |
| Lance Barrett | 8,184 | 0.2989 | 0.2944 |
| CB Bucknor | 7,915 | 0.2972 | 0.293 |
| Phil Cuzzi | 8,086 | 0.2968 | 0.2928 |
| Rob Drake | 7,028 | 0.2951 | 0.2911 |
| Malachi Moore | 7,700 | 0.2939 | 0.2905 |
| Jerry Meals | 3,759 | 0.2974 | 0.2905 |
| Brian O'Nora | 7,221 | 0.2932 | 0.2898 |
Scroll horizontally for all columns.
platoon vs lhp
On-base rate vs LHP (descriptive)
pooled_mean 0.3238 · kappa 226.3 · alpha 73.29 · beta 153.01 · n_entities 394
floor: pa_vs_l>=20
Biggest regressors
The small-n rows the raw leaderboard overstated most.
| Name | n | Raw rate | Shrunk rate | Moved by |
|---|---|---|---|---|
| Jesús Sánchez | 105 | 0.1905 | 0.2816 | -0.0911 |
| Nick Ahmed | 138 | 0.1812 | 0.2698 | -0.0886 |
| Gabriel Arias | 144 | 0.1806 | 0.2681 | -0.0876 |
| Brett Baty | 106 | 0.1981 | 0.2837 | -0.0856 |
| Mitch Garver | 148 | 0.4527 | 0.3748 | +0.0779 |
| Sheldon Neuse | 110 | 0.2091 | 0.2863 | -0.0772 |
| JJ Bleday | 115 | 0.2087 | 0.285 | -0.0764 |
| James McCann | 185 | 0.1946 | 0.2657 | -0.0711 |
| Alek Thomas | 177 | 0.2034 | 0.271 | -0.0676 |
| Jackie Bradley Jr. | 117 | 0.2222 | 0.2892 | -0.067 |
| Jazz Chisholm Jr. | 133 | 0.218 | 0.2847 | -0.0666 |
| Connor Wong | 126 | 0.2222 | 0.2875 | -0.0653 |
Scroll horizontally for all columns.
Leaderboard after shrinkage
Ranked by shrunk rate — the honest ranking.
| Name | n | Raw rate | Shrunk rate |
|---|---|---|---|
| Paul Goldschmidt | 308 | 0.4383 | 0.3898 |
| William Contreras | 293 | 0.4369 | 0.3876 |
| Robbie Grossman | 286 | 0.4266 | 0.3812 |
| Yandy Díaz | 284 | 0.4155 | 0.3749 |
| Mitch Garver | 148 | 0.4527 | 0.3748 |
| Mookie Betts | 359 | 0.4011 | 0.3712 |
| Rob Refsnyder | 219 | 0.4201 | 0.3712 |
| Yordan Alvarez | 367 | 0.3978 | 0.3696 |
| Harold Ramírez | 246 | 0.4106 | 0.369 |
| Jose Altuve | 275 | 0.4036 | 0.3676 |
| Chas McCormick | 254 | 0.4016 | 0.365 |
| Xander Bogaerts | 320 | 0.3937 | 0.3648 |
Scroll horizontally for all columns.
Confounds
- shrinkage assumes the entities are exchangeable draws from one prior -- a modeling choice, not a fact;
- the OOZ rate is descriptive, NOT a called-strike/framing skill and NOT predictive;
- the raw vs shrunk gap is the point: it shows how much a small-sample rate should be discounted, it is not a new 'true skill' claim;
- fixed 2022-2023 window.
Descriptive statistical exhibit only. The shrunk rate is a regularized estimate under a modeling assumption, not a validated forecast-quality finding.