AL-90 · Selected Studies — Sport Management

Umpire Accuracy Lab

MLB umpires called 34,294 incorrect ball-strike calls during the 2018 season — an average of 14 per game, 1.6 per inning. Worse: the error rate rises significantly with two-strike counts, where umpires are biased toward calling false strikes. This lab explores those findings (inspired by the BU Questrom research program on MLB umpire accuracy) in two modes: a measurement lesson on error rates and bias detection, and a structured debate on automated ball-strike technology.

📊 Data: Inspired by: BU Questrom umpire accuracy research program (Mark T. Williams + graduate students); 11-season dataset analysis, ~4 million pitches. See also: bu.edu/articles/2019/mlb-umpires-strike-zone-accuracy/
🔒All inputs are processed locally in your browser. Nothing is transmitted, stored, or logged. Inputs disappear when you close the tab.
Section A — Scale of the Problem
Incorrect calls (2018)
34,294
Ball-strike calls, full season
Per Game
~14
Incorrect calls per game
Per Inning
~1.6
Incorrect calls per inning
Games ended wrong (2018)
20
Final call of game was incorrect
11-season dataset
~4M
Pitches analyzed (inspired by BU)
Games ended wrong (11 seasons)
55
Final call was incorrect
Umpire accuracy rate (est.)
~85%
Borderline calls only (est.)
ABS accuracy rate (est.)
>99%
Hawk-Eye / TrackMan systems
Note on these figures: The 34,294 and associated per-game/per-inning figures are inspired by and consistent with findings from the BU Questrom research program on umpire accuracy. Exact values may vary by study parameters and season analyzed. See bu.edu/articles/2019/mlb-umpires-strike-zone-accuracy/ for published context.
Section B — Two-Strike Count Bias
The Pattern: With a 2-strike count (0-2, 1-2, 2-2, 3-2), umpires are psychologically predisposed to "protect the hitter" — they become less likely to call strike 3 on borderline pitches. This results in elevated false-ball rates on pitches that are actually strikes when the batter is one strike away from striking out.

This is consistent with a broad body of cognitive psychology research on loss aversion under pressure: calling strike 3 on a questionable pitch feels like a decisive, consequential act. The umpire hesitates. The borderline pitch becomes a ball. The pitcher who deserved a strikeout loses it.
Illustrative Error Rates by Count Type

Illustrative rates based on direction and magnitude of BU research findings, not exact study percentages.

No-strike counts (lower error rate)
0-0
~4%
1-0
~4%
2-0
~4.5%
3-0
~4.4%
One-strike counts (moderate error rate)
0-1
~5.5%
1-1
~5.8%
2-1
~5.7%
3-1
~5.5%
Two-strike counts (elevated error rate — two-strike bias)
0-2
~8%
1-2
~8.5%
2-2
~9%
3-2
~9.5%
Section C — Performance Evaluation Framework for Umpires
MLB umpires are one of the most measurable professional workers in the world. Every decision they make is logged, timestamped, and cross-referenceable against a ground-truth measurement system. Most workers — surgeons, teachers, lawyers — face nothing comparable.
ACCURACY RATE
Correct calls ÷ Total calls. Overall accuracy across all umpires typically 85–90% (borderline pitches only). Clear balls and obvious strikes are nearly always called correctly. The ~10–15% error rate applies specifically to pitches near the strike zone boundary.
BIAS COEFFICIENT
Asymmetry between missed balls (calling a ball a strike) and missed strikes (calling a strike a ball). Two-strike count analysis reveals a systematic bias toward missed strikes — umpires are more likely to err toward keeping batters alive. This is not random noise; it is directional, predictable error.
COUNT-SPECIFIC ACCURACY
Performance degrades under specific high-pressure conditions — a critical performance evaluation insight for any worker. A surgeon who performs 95% of surgeries well but degrades under time pressure or specific patient profiles has a very different risk profile than one who performs uniformly. Umpire accuracy falling in 2-strike counts is a conditional performance problem.
IMPLICATIONS FOR PERFORMANCE MANAGEMENT
If 99% of an umpire's calls are correct on obvious pitches but they struggle on borderline 2-strike counts, a performance improvement intervention should target specifically that condition — not overall retraining. The data enables targeted interventions impossible without granular performance logging.
Discussion Questions
  • If you were an MLB team, how would you use umpire accuracy data? (e.g., game preparation, at-bat strategy, pitcher selection for high-stakes counts)
  • Is an accuracy rate of ~85% acceptable for a professional whose errors directly affect game outcomes and potentially team playoff qualification?
  • What does "two-strike count bias" suggest about cognitive load and decision-making under pressure? What other professions show similar conditional performance degradation?
  • If umpire error is not random but systematic, who bears the cost? Are pitchers harmed more than batters, or vice versa?
  • Should individual umpire accuracy data be publicly disclosed? What are the arguments for and against transparency?
Automated Ball-Strike (ABS) Technology — Structured Debate

ABS uses camera-tracking systems (Hawk-Eye, TrackMan) to call balls and strikes automatically. It is currently used in Triple-A and independent leagues. Select a position to view a full argument set.

Error Analysis Framework — For Any Performance Evaluation System
The umpire accuracy question is a special case of a general problem: how do you evaluate whether a decision-making system (human or automated) is performing adequately? Five questions structure any rigorous analysis.
1
What is the reference standard?
An "error" is only definable if there is a ground truth to compare against. For umpires: does the pitch cross the geometrically defined strike zone? Hawk-Eye and TrackMan provide that ground truth. Without a reference standard, you cannot measure accuracy — you can only measure consistency. Critical implication: Many real-world performance evaluations lack a genuine reference standard (what is the "correct" legal ruling? the "correct" diagnosis for a complex case?). Umpiring is unusually tractable because the ground truth is physical and measurable.
2
What is the capture rate?
Are all decisions logged? For umpires: yes — every pitch is tracked by MLB's Statcast system. This is a nearly unique situation in professional performance evaluation. Most workers' decisions are not logged at a 99%+ capture rate. Teachers' moment-to-moment instructional decisions are not captured. Surgeons' judgment calls in the operating room are not captured. The high capture rate is what makes umpire performance evaluation possible at this level of precision. Key question: What is the capture rate for the decisions you care about evaluating?
3
Are errors random or systematic?
Random noise is unavoidable and theoretically unbiased — it doesn't favor any party over time. Systematic error (bias) is directional, persistent, and correctable. The two-strike count bias in umpire calling is systematic: it favors batters over pitchers in high-leverage counts, predictably and consistently. A random error rate of 10% is fundamentally different from a systematic bias rate of 10%. Why this matters: Systematic errors can be corrected through training, incentives, or technology. Random errors can only be reduced, not eliminated. Knowing which you're dealing with determines your intervention strategy.
4
What are the consequence asymmetries?
Is a false strike (calling a ball a strike) worse than a false ball (calling a strike a ball)? In different contexts: with a batter ahead 3-0, a missed strike matters more than a missed ball. In a 0-2 count, a missed third strike costs a pitcher an out. ABS debates partly rest on this: if errors were symmetric, the efficiency argument for ABS would be weaker. The two-strike bias suggests umpire errors are asymmetric — they disproportionately help batters at the margin that matters most. General principle: All error analysis should quantify consequence asymmetries, not just error rates.
5
What is the comparison benchmark?
How good could this process be? The technology limit for ABS is known: ~99.5% accuracy. The human limit is less certain, but consistent evidence suggests 90–92% on borderline pitches is near the ceiling for a human umpire working at game speed. Without a benchmark, "85% accuracy" sounds like a number without context. With ABS as a comparison, it becomes clear that ~10 percentage points of accuracy are available through technology adoption — the question is whether that gain is worth the costs of transition. Ask always: What is the best this system could theoretically be? Are we near that limit, or far from it?
Apply This Framework to Your Industry

The five-question framework applies to any human decision-making system subject to error analysis. Consider these applications:

Medical Diagnosis
Reference standard: Biopsy, lab result, confirmed outcome. Capture rate: Varies — some decisions are logged, many are not. Systematic errors: Diagnosis bias by patient demographics (documented in multiple studies). Consequence asymmetry: False negative (missing a disease) typically worse than false positive. Benchmark: Computer-aided diagnosis (CAD) for mammography reads, AI for radiology — what is their measured accuracy vs. human radiologists?
Legal Decisions (Sentencing)
Reference standard: Hard to define — what is the "correct" sentence? Capture rate: Sentencing decisions are logged. Systematic errors: Documented disparities by defendant race, time of day (Israeli parole board hunger study). Consequence asymmetry: False imprisonment vs. wrongful release — deeply asymmetric consequences. Benchmark: Actuarial risk tools vs. judicial discretion — ongoing empirical debate.
Financial Audit
Reference standard: Actual financial statement accuracy (knowable after the fact, sometimes). Capture rate: Auditor decisions are documented in workpapers. Systematic errors: Auditors who develop long relationships with client management may develop confirmation bias. Consequence asymmetry: Missing a fraud (false negative) vs. flagging a non-issue (false positive) — different regulatory and reputational consequences. Benchmark: AI-assisted audit tools are now being evaluated against human auditor accuracy in large-sample studies.