How an AI Math Proof Reveals Law's Verification Crisis
The same AI that solved an 80-year-old math conjecture also produces confident-sounding errors half the time — a verification failure pattern that directly maps to the hallucination crisis now generating 1,000+ US sanction rulings. This article examines the Leiden Declaration's warning and what it means for attorneys relying on AI tools.
- Jurisdiction
- US Federal
- Court
- United States District Courts
- AI tool named
- OpenAI o4-mini
- Ruling date
- Jun 2, 2026
- Source document
- View primary court order ↗
- Last verified
- Jul 25, 2026
Lex Machina Review is an independent risk-tracking and reference resource. Nothing on this site is legal advice, and using it does not create an attorney-client relationship. Every record is reviewed against primary sources but may not reflect the most current status of a matter — always verify directly against the cited court order, rule text, or a licensed attorney before relying on it.
Companion explanation — secondary to the source document above
For anyone trying to understand the OpenAI unit-distance story, the shortest honest answer is also the most uncomfortable one: the reported result may be a genuine mathematical achievement, and it still does not prove that AI reasoning is dependable in the way lawyers, courts, or clients need it to be.
OpenAI’s model reportedly produced a disproof of the Erdős unit distance conjecture, a problem that had resisted mathematicians for roughly 80 years. That is not a trivial headline. A machine reaching into a long-stalled mathematical problem is exactly the kind of technical event that deserves attention. But the same reporting also contained the fact that should slow every professional reader down: OpenAI researcher Sébastien Bubeck said the model produced the correct disproof in only 50% of trials at maximum token budget, and the incorrect attempts were not released for outside inspection.[1]

That is the legal-risk version of the story. Not whether the system can ever be brilliant. It apparently can. The question is what happens to the person who receives a particular output and cannot tell whether this is the brilliant half or the false half.
The Breakthrough Arrived With Its Own Verification Problem
The 50% figure should not be treated as a settled benchmark for all future AI math systems. The reported conditions remain narrow: the number of trials, the exact model configuration, and the unreleased failure cases are not available in a form that lets outsiders reproduce the claim. That matters. A failure rate disclosed through reporting is not the same thing as a peer-reviewed reliability study.
Still, the number is load-bearing because it exposes the asymmetry. The public sees the successful proof. The community does not get the corresponding wrong proofs. Melanie Matchett Wood of Harvard said mathematicians need to know how often the system “produced an incorrect solution with flawed reasoning,” and that is not a request for gossip about failed drafts. It is a request for the evidence needed to decide whether the system’s confidence has any reliable relationship to correctness.[2]
A vendor announcement can celebrate the one successful run. A working mathematician, judge, associate, or clerk has to deal with the next run. That next output does not arrive stamped with its true status. It arrives in polished prose, with internal structure, citations or lemmas or case names, and a surface that encourages the human reader to relax.
Terry Tao put the problem sharply: AI is “much better at sounding like they have the right answer than actually getting it… right or wrong, they will always look convincing.” Ken Ono, after seeing OpenAI’s o4-mini in a 2025 secret meeting, said the model had “mastered proof by intimidation.”[3] That phrase is useful because it names a professional hazard rather than a personality flaw. The output does not have to persuade by evidence. It can persuade by looking too complete to doubt.
Why “Proof by Intimidation” Sounds Familiar in Court
Legal AI failures already follow the same pattern. The dangerous filing is rarely a page of obvious nonsense. It is formatted like a brief, written in ordinary legal cadence, and often confident enough that a tired reviewer can mistake coherence for verification.
The analogy has limits. A mathematical proof, at least in principle, can sometimes be pushed into formal verification systems that check whether each step follows from accepted rules. Legal work generally cannot be reduced to that kind of proof object. A citation can be checked; a quotation can be compared to the source; a procedural rule can be read. But legal judgment also depends on jurisdiction, posture, adverse authority, facts, timing, remedies, and professional duties. The court does not sanction a model. It sanctions, questions, or distrusts the human professionals who filed the work.
That difference makes the legal setting less forgiving, not more. If a math model produces a false proof, the community may reject it after review. If a legal model produces a false authority and a lawyer files it, the harm can move through a client matter, an opposing party’s response, a clerk’s workload, and a judge’s docket before the failure is fully exposed.

The Leiden Declaration Reads Like a Risk Taxonomy
On June 2, 2026, the Leiden Declaration gathered more than 3,125 signatories and was endorsed by the International Mathematical Union. Its lead authors include Terence Tao, Peter Scholze, Ursula Martin, and Kevin Buzzard.[4] It is framed around mathematics, but its core warnings describe a class of AI failure that legal professionals already recognize.
The declaration warns that AI systems can produce “plausible but unreliable (or even incorrect) arguments which are difficult to distinguish from correct mathematical proofs.” It also states that “the responsibility for the correctness and adequacy of the arguments and results… remains exclusively with the human authors.”[4] Substitute “brief,” “memo,” “contract analysis,” or “research answer” for “proof,” and the operational duty is not hard to see.
| Leiden warning | Legal risk translation |
|---|---|
| Unreliable arguments that resemble correct ones | Fabricated or distorted authorities can appear in professional legal prose before anyone checks the source. |
| Attribution failure | A model may surface ideas, structures, or analysis without preserving the human lineage needed for professional credit and source evaluation. |
| Hype-driven evaluation | Press releases and demos can shape expectations before courts, firms, or clients see reproducible benchmark data. |
| Incentive distortion | Systems optimized for fluent answers can reward the appearance of completion over verifiable accuracy. |
| Loss of autonomous human judgment | Reviewers may defer to polished output when they lack time, expertise, or protocols to inspect it. |
The first warning is the most direct. A wrong answer that looks wrong is manageable. A wrong answer that looks right consumes the reviewer’s time, transfers the detection burden downstream, and creates a record that may already have influenced someone’s decision before correction arrives.
The second warning, attribution failure, is not just academic etiquette. Wood criticized the OpenAI proof for failing to credit prior ideas and said analogous behavior would be “professional malpractice” if done by a human.[2] In law, attribution is part of verification. Knowing where a proposition came from tells the reader whether it is binding authority, persuasive authority, dicta, a party’s characterization, a treatise summary, or a model’s unsourced reconstruction.
The third warning is about evaluation by publicity. A publicized success can be real and still be insufficient. Lawyers would not accept a litigation support vendor’s accuracy claim merely because it found one important document in one impressive demonstration. The same caution should apply when a frontier model solves a famous problem while the failed attempts remain unavailable.
The fourth and fifth warnings are linked. If a system is rewarded for producing fluent, complete-looking work, and the human reviewer is rewarded for speed, the review process starts to drift. The lawyer becomes less an author and more a passive acceptor of output. That is the point at which “AI assistance” stops being a productivity tool and starts becoming an undocumented delegation of professional judgment.
The Legal Evidence Is No Longer Hypothetical
The legal profession does not need to imagine what plausible but false AI output looks like. Documented AI hallucination cases now include more than 1,490 matters globally, with more than 1,000 U.S. sanction rulings, and Q1 2026 sanctions exceeding $145,000.[5] Those numbers do not prove that every legal AI use is reckless. They do show that ordinary human review has not reliably caught fabricated or defective AI-supported legal work before it reached courts.
Benchmark evidence points in the same direction, though it should be read carefully. A 2025 Stanford RegLab and Stanford HAI benchmark found that legal AI models, including Lexis+ AI and Ask Practical Law AI, hallucinated in more than 17% of benchmark queries.[6] That result is not a permanent measurement of every product’s Q3 2026 performance. Vendors may have improved systems since then; individual use cases vary; benchmark construction matters. But a one-in-six-plus hallucination baseline in legal research tools is not compatible with casual reliance.
The important comparison is not that math and law fail in identical technical ways. They do not. The comparison is that both domains expose a reviewer-facing problem: the model’s answer may look authoritative before its reasoning, sources, and failure modes are inspectable. In mathematics, a small circle of experts may eventually test the proof. In court, the test may arrive as an order to show cause.
This is also why vendor claims about improved refusal behavior deserve caution. A model that says “I cannot solve this” more often would be useful. But without released data showing when it refuses, when it fabricates, and how reviewers can distinguish the two, the claim remains a product assertion rather than a control.
Reliability Is Not Established by the Best Output
Professionals do not buy reliability by pointing to a system’s most impressive answer. They establish it by understanding the distribution of answers: what succeeds, what fails, how often failures occur, whether failures are detectable, and what the human must do before the work can be used.
The OpenAI unit-distance result is therefore both impressive and incomplete as evidence. The correct proof matters. The unreleased incorrect proofs may matter more for anyone trying to decide whether similar systems can be trusted outside a research showcase. If the failed outputs were obviously broken, that would support one operational conclusion. If they were elegant, long, and wrong in ways only specialists could detect, that would support another. Without them, the outside observer is asked to admire the success while guessing at the risk.
Law firms and legal departments should recognize the same pattern in product demonstrations. A tool may summarize one contract beautifully, identify one regulation accurately, or generate one useful first draft. That does not answer the question a supervising lawyer has to answer before relying on it: what does this system do when it is wrong, and how will we know before someone else does?
That question is not solved by better prompt writing alone. Prompt discipline may reduce some errors, and structured workflows can make review easier. But the duty still lands on source-level verification: open the case, read the rule, inspect the quoted language, check jurisdiction and procedural posture, compare adverse authority, and decide whether the generated analysis survives contact with the materials. For lawyers building that habit into junior workflows, the issue is close to the one discussed in AI verification as a career skill for Gen Z lawyers: verification is not clerical cleanup after AI use; it is the professional act that turns output into work product.
What Should Count Before Lawyers Rely on AI Reasoning
The practical standard is not “never use AI.” That would be a poor reading of the math story. The unit-distance result suggests that frontier systems can produce work worth serious expert attention. The practical standard is that legal organizations should not treat frontier capability as risk reduction unless the relevant failure cases, benchmark methods, and verification pathways are independently inspectable.
- A legal AI claim should identify the tested task, not just advertise general reasoning ability.
- Benchmarks should disclose the questions, scoring method, version tested, and treatment of partial or evasive answers.
- Failure examples should be available to qualified reviewers, because the shape of the error determines the control needed.
- Outputs used in legal work should preserve citations, quotations, and source paths in a form a human can inspect.
- Responsibility should remain assigned to a human author who has actually verified the parts being used.
This is where the cost question becomes harder than many AI pilots admit. If a tool saves drafting time but increases source-checking time, the net benefit depends on the workflow, the stakes, and the reviewer’s expertise. Hidden verification time is still professional labor, a point that also appears in discussions of why OpenAI API pricing can mislead legal professionals. The cheap part is generating fluent text. The expensive part is deciding whether it is safe to use.
Fast-moving legal questions make the problem worse. When statutes, agency positions, or litigation postures are unsettled, a model’s confident synthesis can be especially hard to evaluate because there may be no single settled answer to retrieve. That is why AI legal research risk in disputes like the SAVE Act is not merely a hallucination problem; it is a problem of authority, timing, and ambiguity.
The OpenAI math proof should be remembered for both halves of its lesson. A model may reach a result that experts take seriously. The same model, in the same reported setting, may also produce wrong answers often enough that undisclosed failures become central to any reliability judgment. Until the legal profession can inspect failure cases, benchmark methods, and source-level evidence, a confident AI answer is not a legal work product. It is an item awaiting verification by a responsible human author.
References
- An AI math breakthrough sparks calls for new guardrails, ScienceNews, June 8, 2026.
- OpenAI announces AI’s biggest math breakthrough yet, Scientific American.
- Proof by intimidation: AI is confidently solving impossible math problems, Live Science.
- The Leiden Declaration, leidendeclaration.ai, June 2, 2026.
- AI Hallucination Cases Database, Damien Charlotin.
- AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries, Stanford HAI, 2025.
Related records
Tool profile
Browse tool evaluations →Governing regulation
The 2025 DACA Protection Bills, Provision by ProvisionPreventive workflow
Browse verification workflows →
Report a correction or tip
Spotted an outdated figure, a misstated fact, or a ruling this case record should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.
Report a correction or tip for this record →