Skip to main content
What Legal AI Can Learn from Baseball's AI Challenge System
market dataSource type: trade publication

What Legal AI Can Learn from Baseball's AI Challenge System

The MLB's Automated Ball-Strike (ABS) challenge system, adopted after years of testing, offers a replicable governance model for law firms deploying AI tools: preserve human judgment as the primary decision-maker while using AI for high-leverage verification. This article examines how the challenge system's phased rollout, stakeholder feedback loops, and transparency safeguards transfer to legal AI workflows in document review, contract analysis, and e-discovery.

Updated

Major League Baseball did not bring Statcast-powered ball-strike tracking into the regular season by removing the umpire from the plate. It chose a narrower, more defensible design: the human calls the pitch, a player may challenge in a tightly limited window, and the machine verifies only the disputed moment. That choice is the part legal teams should study.

The point is not that a strike call and a privilege determination are the same kind of judgment. They are not. The point is that MLB faced a familiar governance problem: a professional standard that appears crisp from a distance becomes contested when applied at speed, by humans, under consequence. Cornell researchers described the strike zone as a “social construct” marked by expert disagreement, and warned that the Automated Ball-Strike system effectively “changed the rules so that it would be accepted” by translating the three-dimensional strike zone into a two-dimensional plane at the front of home plate.[1]

That is the legal AI problem in miniature. A contract-risk classifier, e-discovery relevance model, or privilege screen can be sold as verification. In practice, it may also redefine what the organization treats as risky, responsive, privileged, standard, or worth escalating. If the workflow does not make that change visible, the clean dashboard becomes a bad record.

Baseball umpire and legal professional connected by a glowing bridge

The important design choice was not AI. It was who retained authority.

By the time MLB approved the ABS Challenge System for 2026, human umpires were not performing badly. MLB reported that umpire accuracy had improved from 84.1% in 2008 to 92.8% in 2025.[2] That matters because it removes the laziest version of the AI argument. This was not a rescue operation for an incompetent profession.

And yet, during Spring Training, 52.2% of challenged calls were overturned.[2] Put those numbers together carefully. They do not prove that machines should call every pitch. They show that even a high-performing expert can benefit from a second look when the affected person identifies a high-leverage call as worth reviewing.

That is a better model for legal work than the replacement story. In document review, the useful question is not whether AI can replace reviewer judgment across the entire corpus. It is whether the system can identify the slices where a second look is most valuable: inconsistent coding, borderline privilege calls, unusual clause language, missing exhibits, or a production set that does not match the review protocol. The professional still owns the decision. The machine makes the dispute easier to see.

The distinction is practical. If a law firm deploys an AI tool to auto-classify every indemnity clause and silently route “low risk” contracts past counsel, it has moved judgment into the model. If the same tool flags clauses that diverge from an approved playbook and requires counsel or a trained contract manager to accept, revise, or override the flag, it has created a reviewable control. The difference will matter later, when someone asks why the contract was approved.

MLB took seven years to make the workflow boring enough to trust

The most useful part of the ABS story is the rollout, not the hardware. MLB tested the system in the Atlantic League in 2019, then in Triple-A from 2023 through 2025, before adopting the challenge format in MLB for 2026.[2] For legal operations leaders, that sequence is more valuable than any demonstration video. It shows a technology moving through progressively higher-stakes environments while the operating rules were still open to revision.

Three-level staircase with feedback arrows showing phased rollout and iteration

The league also surveyed the people who had to live with the system. MLB reported 72% fan approval of the challenge system, and survey results showed 60% of Triple-A players and coaches preferred the challenge format, compared with 24% who preferred full ABS and 16% who preferred human-only calls.[3][4] That split is the governance signal. The affected stakeholders did not reject machine verification. They preferred it with human judgment still in the primary position.

Legal AI pilots often skip the comparable step. A procurement team sees a compelling demo, a practice group chair wants faster throughput, and the pilot measures time saved. Those are not irrelevant metrics, but they are incomplete. The reviewers who must correct the output, the attorney who signs the filing, the information-governance team that must defend the process, and the compliance officer who must explain the audit trail all need a voice before the workflow becomes standard operating procedure.

A legal rollout modeled on ABS would not begin with enterprise-wide adoption. It would start with a bounded workflow where the stakes are real but contained. It would document which decisions remain human, which outputs are advisory, which disputes trigger escalation, and which failure modes require reverting to the prior process. Then it would ask the people closest to the work whether the tool helped them exercise judgment or merely gave them new cleanup duties.

ABS governance choiceLegal AI analogue
Tested first in lower-stakes environments before MLB adoptionPilot in a bounded matter type, contract category, or review population before broad rollout
Challenge format preferred over full automation by surveyed Triple-A players and coachesUse AI to surface, verify, and triage disputed issues while keeping professional sign-off primary
Limited number of challenges and short challenge windowDefine when users may invoke AI review, who may override, and how long unresolved exceptions can sit
MLB-controlled review and fallback protocolsCentralize audit trails, privilege protections, access controls, and reversion rules when confidence fails

This is also where legal teams should be careful with cost-efficiency narratives. Faster inference, cheaper compute, and lower per-document review costs are worth tracking, but they do not answer the professional-responsibility question by themselves. A lower-cost model that produces more unreviewed assumptions may be expensive in the only place that matters: the record. That same verification gap appears in other AI procurement settings, including debates over Gemini and legal AI cost efficiency, where savings need to be tied to defensible use rather than treated as the whole business case.

The safeguards are the lesson

ABS is not just a tracking system. It is a rule-bound challenge process. Only players can initiate challenges. A challenge must be made within two seconds. Each team receives two challenges per game. MLB reported that 71% of fans wanted four or fewer total challenges, and the average review took 13.8 seconds.[3][4]

Those constraints do real work. A limited challenge count discourages reflexive review. A short challenge window prevents teams from waiting for outside analysis before deciding. Player-only initiation keeps the decision close to the person who saw the pitch, instead of moving it to a remote command center with different incentives. The short review time also matters because a verification system that constantly stops the underlying work eventually becomes the work.

Translate that to legal workflows and the parallels are immediate. A contract team can require that AI escalation be invoked by the lawyer or contract manager responsible for the document, not by an unrelated analytics team chasing variance. An e-discovery team can limit AI recoding requests to defined quality-control events rather than allowing every disagreement to reopen the review. A compliance workflow can impose a short decision period for accepting, rejecting, or escalating a model recommendation so exceptions do not disappear into a queue.

The point is not to make legal work imitate baseball’s timing rules. A two-second window would be absurd in a privilege review. The transferable principle is that a challenge mechanism needs boundaries before the dispute arises. If the user can invoke AI review whenever the result is convenient, or ignore it whenever the result is inconvenient, the organization has not designed a control. It has purchased optional cover.

MLB also built controls around information timing and data ownership. The Gameday App data is delayed, the broadcast strike-zone box is delayed, team-owned tracking systems are prohibited for challenge purposes, and MLB reviews challenge video.[3] Those are not decorative details. They protect the integrity of the challenge decision by limiting who can see what, when, and for what purpose.

Legal AI needs the same seriousness about information asymmetry. A privilege tool that exposes too much metadata to the wrong team can create a waiver problem. A contract-analysis tool that lets business users see model confidence without context may encourage them to pressure counsel into accepting an output that should be reviewed. A litigation analytics dashboard that permits after-the-fact edits without audit history is an invitation to lose the story of how the decision was made.

Good controls are often unglamorous: role-based access, preserved prompts and outputs where appropriate, matter-level configuration logs, mandatory human sign-off, override reasons, privilege-aware storage, and clear rules for when model output cannot be used. They are also what make an AI workflow explainable to a court, regulator, client, insurer, or internal audit function. For firms comparing AI governance duties across jurisdictions, those controls should sit alongside the broader compliance mapping discussed in state AI law compliance for law firms.

Fallback rules matter before the failure occurs

The system also had to answer a simple question that many AI deployments avoid until too late: what happens when tracking fails? MLB reported only four untrackable pitches out of 88,534 in Spring Training, a 0.0045% failure rate, with defined “call stands” fallback protocols.[2] The small number is reassuring. The fallback rule is more important.

Legal teams should write the equivalent rule before the first production deadline. If an AI privilege screen is unavailable, does the team revert to manual review, sample-based quality control, or a prior approved tool? If a model cannot process a file type, is the document excluded, converted, separately reviewed, or escalated? If a contract tool gives no recommendation, does silence mean acceptable, unreviewed, or blocked?

The most dangerous fallback is the one no one writes down: people improvise under deadline, the workaround becomes invisible, and later the organization describes a process that is cleaner than the one actually used. Anyone who has reconstructed a messy review record knows that this is where “efficiency gains” become expensive.

A legal version of the ABS challenge model does not require turning every task into a formal appeal. It requires separating primary judgment from machine verification and documenting how the two interact.

  • In contract review, the human reviewer applies the playbook first. AI then flags provisions that appear inconsistent with approved language, missing from the draft, or unusually allocated. The reviewer accepts, rejects, or escalates the flag with a recorded reason.
  • In e-discovery, reviewers code documents under the review protocol. AI surfaces clusters of inconsistent coding, probable privilege conflicts, or documents similar to known hot documents. A senior reviewer or attorney resolves the disputed population.
  • In compliance monitoring, the business process continues under existing controls. AI identifies transactions or communications that match defined risk indicators. Compliance personnel decide whether the alert is closed, investigated, or escalated.
  • In legal research, the lawyer develops the issue and reviews the authorities. AI can check for omitted adverse authority, citation problems, or conflicts among propositions. The lawyer remains responsible for the filing or advice.

This structure is less glamorous than a promise to automate the workflow. It is also more likely to survive contact with professional obligations. The lawyer or legal professional closest to the decision is not reduced to a rubber stamp. The AI output is not treated as neutral merely because it is mathematical. The organization can show where the tool was used, what it saw, who acted on it, and what happened when the human disagreed.

Outages and model failures make this distinction sharper. If the workflow depends on uninterrupted model availability, the fallback process must be real, tested, and known to the people doing the work. The professional-responsibility implications are similar to those raised by Claude AI outages and legal risk: the vendor’s problem can become the lawyer’s problem if the firm has no defensible backup.

The analogy breaks where law is least like a strike zone

The caution comes from the same baseball evidence that makes the governance model useful. Cornell’s point was not simply that baseball used technology. It was that the technology made acceptance easier by changing the applied standard: a three-dimensional strike zone became a two-dimensional plane.[1]

Three-dimensional cube compressed into a flat plane while professionals observe

That is manageable in baseball because the league can decide what version of the rule it wants to administer. Law is less forgiving. Legal standards are open-textured, fact-dependent, and often contested by design. “Reasonable efforts,” “materiality,” “good cause,” “ordinary course,” “proportionality,” and “commercially reasonable” are not physical locations waiting to be measured. They are legal judgments made within institutional, factual, and adversarial contexts.

So a legal AI tool that classifies clauses does not merely find the “real” category. It may train the organization to treat certain clause forms as standard and others as deviant. A litigation-risk model may make some facts look legally salient because they are easy to encode, while pushing harder-to-measure facts out of the decision path. A research tool may privilege authorities that fit its retrieval structure and make less conventional arguments harder to see.

This is where vendor claims about replacement deserve particular skepticism. If a system says it can automate a legal standard, the first question should be what it had to flatten in order to do so. The answer may be acceptable for a narrow triage task. It may be unacceptable for a dispositive legal judgment. The difference should be decided by lawyers and accountable stakeholders before deployment, not discovered in a sanctions motion, client dispute, regulatory inquiry, or post-closing loss review.

The same caution applies to tools that look capable but create ethics traps through hidden assumptions, unclear data handling, or unverifiable outputs. That is why model evaluation should include not only performance and cost, but also the professional-use questions raised in analyses such as the Kimi K3 law firm ethics trap.

Use AI to sharpen judgment, not to hide where judgment moved

MLB’s ABS Challenge System is useful to legal teams because it did not pretend that a contested professional judgment became simple once a machine entered the room. It preserved the umpire’s call, gave affected participants a limited way to challenge, made verification fast and visible, controlled the surrounding information, and wrote a fallback rule for failure.

That is the better pattern for legal AI: phased deployment, stakeholder input, limited override mechanisms, transparent verification, and mandatory human responsibility for the final act. The unresolved part must stay visible too. In law, the categories are not fixed physical planes. A workflow that uses AI to verify judgment must be designed carefully enough that it does not silently rewrite the rule it claims only to apply.

References

  1. AI on deck: assessing impact of MLB's new ball-strike system, Cornell Chronicle, March 2026.
  2. MLB to use ABS Challenge System starting in 2026, MLB.com.
  3. MLB ABS System Explainer, MLB.com.
  4. MLB AI-powered challenge system in 2026, AI Business.

Corrections & feedback

Submit corrections, flag outdated information, or provide additional market context. Comments are moderated.

Comments

Join the discussion with an anonymous comment.

Loading comments...
Blogarama - Blog Directory