Why OpenAI Astra's Math Isn't Enough for Legal Tech
OpenAI Astra's Lean-verified math proofs don't transfer to filing-safe legal arithmetic. What litigators should weigh instead is unverified error-rate data on real damages, interest, and fee calculations — where documented hallucination and multi-step failure rates, not proof milestones, are the operative risk signals.
- Jurisdiction
- United States
- Court
- Various U.S. federal and state courts
- AI tool named
- OpenAI Astra
- Ruling date
- Aug 1, 2026
- Source document
- View primary court order ↗
- Last verified
- Aug 3, 2026
Lex Machina Review is an independent risk-tracking and reference resource. Nothing on this site is legal advice, and using it does not create an attorney-client relationship. Every record is reviewed against primary sources but may not reflect the most current status of a matter — always verify directly against the cited court order, rule text, or a licensed attorney before relying on it.
Companion explanation — secondary to the source document above
OpenAI’s Astra math result is the kind of announcement legal tech teams should read carefully, not dismiss. OpenAI reported on August 1 that its system produced ten advances in mathematics and theoretical computer science, with proofs formalized in Lean; reporting around the release identified the internal model as Astra.[1][2] For legal tech teams, though, the decisive question is narrower: does that achievement show that the model can calculate damages, prejudgment interest, deadlines, or fee awards safely enough to put the number in a filing?
On the public evidence now available, no. Astra is an internal, unreleased model, and there is no independent benchmark showing how it performs on real legal arithmetic: messy facts, rate changes, partial payments, date counting, fee entries, local rules, and a lawyer’s signature block at the end. The proof milestone is important. It is not a filing-safety credential.

What Astra actually proved
The reported facts are genuinely striking. OpenAI said the system contributed to ten previously unsolved problems, with Lean-verified certificates, and reported FrontierMath performance of 89% on Tier 1–3 problems and 83% on Tier 4 problems. OpenAI also described the compute cost as roughly $2,000 and noted human-assisted preparation of manuscripts around the results.[1] Those are not casual chatbot answers dressed up as mathematics. A Lean certificate means the proof has been checked inside a formal system against explicit rules.
That matters because theorem proving is a domain where verification can be built into the output. The system does not merely say, “trust me”; the proof can be mechanically checked. If the formalization is correct and the checker accepts it, the result has a kind of auditability that ordinary natural-language AI output lacks.
Legal arithmetic rarely arrives in that shape. A damages calculation begins with records that may be incomplete or disputed. A prejudgment-interest number may require selecting the right statutory or contractual rate, applying it over changing periods, deciding whether compounding is allowed, and stopping or restarting the clock after partial payments. A fee petition depends on time entries, reductions, billing judgment, rates, tasks, staffing, and local practice. Date calculations can turn on service method, weekends, holidays, emergency orders, and jurisdiction-specific rules. The dangerous part is often not long division. It is choosing the right inputs and rule path before the arithmetic begins.

The benchmark mismatch is the point
Released reasoning models still show error patterns that matter for litigation. OpenAI’s own PersonQA testing, as reported by TechCrunch, found higher hallucination rates for newer released reasoning models than for o1: o3 at 33% and o4-mini at 48%, compared with o1 at 16%.[3] PersonQA is not a damages benchmark. It is a factual-recall benchmark. But it is a useful warning against treating “reasoning model” as a synonym for “reliable factual assistant.”
Legal-specific research is not more comforting. Stanford HAI reported that legal AI research tools hallucinated in 17% to 34% of benchmarking queries, while general GPT-4 responses contained errors on 58% to 82% of legal queries.[4] Those figures do not measure Astra, and they do not measure every legal task. They do show that legal-domain presentation can coexist with legally material false output. A tool can look fluent, cite authority, and still leave the lawyer with a verification problem.
Arithmetic has its own failure mode. Dojo Labs, in a vendor-sourced audit that should be treated as directional rather than independent, reported that LLMs without external computation tools fail multi-step arithmetic roughly 40% of the time and that a frontier GPT-5 system produced wrong final answers on five-step compound-interest problems 58% of the time.[5] The caveat matters: Dojo sells audit services, and its figures should not be converted into a universal industry rate. Still, the task type is closer to legal damages work than a Lean-certified theorem is. Compound-interest problems, unlike abstract proof certificates, ask the system to carry intermediate quantities through a sequence of operations without drifting.
This is why a single impressive benchmark cannot do the work procurement teams want it to do. A 2026 PNAS analysis titled “There is no free benchmark” makes the institutional point directly: benchmark results are not interchangeable measures of general reliability.[6] FrontierMath, PersonQA, legal-research hallucination tests, and arithmetic audits each test different behaviors under different conditions. None of them, by itself, answers whether an AI system can calculate a statutory-interest award from a real docket record and produce a number a lawyer can sign.
Where legal arithmetic becomes filing risk
The litigation risk is not that an AI model sometimes gets a schoolbook equation wrong. The risk is that a number moves from draft work product into a pleading, declaration, fee motion, settlement demand, or proposed judgment. At that point, someone has represented the calculation to a court, opposing counsel, a client, or an insurer.
| Legal calculation | What the model must do | Where the error becomes material |
|---|---|---|
| Damages totals | Extract amounts, classify categories, avoid double-counting, apply offsets | Demand letters, expert schedules, pleadings, proposed judgments |
| Prejudgment interest | Identify the correct rate, period, compounding rule, start date, stop date, and payment credits | Judgment submissions, settlement evaluations, post-trial motions |
| Deadlines | Count days under the applicable rule, account for weekends, holidays, service method, and local modifications | Filing calendars, default motions, appeal notices, response dates |
| Fee petitions | Sum time entries, apply reductions, distinguish compensable and noncompensable work, calculate lodestar adjustments | Fee motions, declarations, billing audits, settlement fee allocations |
Interest is a particularly bad place to confuse proof capability with filing reliability. A model may handle the arithmetic operation in isolation but invent or misread the rate period. It may apply simple interest where the contract calls for compounding, or compound a subtotal that should have been reduced by an interim payment. In a real filing, the bad output does not announce itself as a hallucination. It appears as a clean dollar figure with a plausible explanation.
Fee calculations create a different exposure. The model may add the time entries correctly but include work excluded by the court’s prior order, collapse attorney and paralegal rates, or miss a voluntary haircut already taken in the declaration. Deadline calculations are harsher still because a single wrong date can convert a manageable correction into a default, waiver, or malpractice problem. These are workflow failures, not abstract math failures.

The ethics frame does not change because the model is better at proofs
ABA Formal Opinion 512, issued in July 2024, treats generative AI tools as nonlawyers for purposes of supervision under Model Rule 5.3, while also tying AI use to the lawyer’s competence duties under Model Rule 1.1.[7] That framing is blunt in the best way: the tool can assist, but it does not absorb the lawyer’s duty to check the work. Rule 11 points in the same practical direction in federal litigation. The signature certifies that a reasonable inquiry has been made. A model’s internal confidence, vendor label, or mathematical pedigree is not the inquiry.
Even favorable legal-practice guidance assumes verification. Debevoise’s Data Blog recommended o3-pro for tasks including “double-checking critical legal arguments or calculations,” but that recommendation is not an invitation to file unreviewed numbers.[8] The useful role is as a second pass, a consistency check, or a prompt to inspect assumptions. The lawyer still has to decide whether the rate table is current, whether the damages schedule uses the right denominator, and whether the final number ties back to admissible records.
That is the same verification problem running through AI filing failures more broadly. The site’s prior analysis of the AI memory bottleneck in legal work treated verification as the safeguard that actually matters under professional-responsibility rules. The same nondelegable duty appears in sanctions coverage involving AI-hallucinated immigration briefing: the duty attaches when the lawyer signs and files, not when the software generates.
Citation sanctions are not math sanctions, but the lesson travels
The reported sanctions cases remain mostly citation cases, and that distinction should be kept clean. EDRM’s 2026 review described an Oregon per-infraction schedule of $500 per fabricated citation and $1,000 per fabricated quotation, reported Q1 2026 AI-sanctions penalties exceeding $145,000, and discussed familiar cases including Mata v. Avianca, Whiting v. City of Athens, Couvrette v. Wisnovsky, and Ghiorso.[9] Mata involved a $5,000 sanction for six fake cases; Whiting involved a $30,000 penalty.[9]
Those cases do not establish a court rule for AI-miscalculated damages. In the materials reviewed here, no reported sanctions decision was found that specifically punished a lawyer for an AI-generated arithmetic error in a legal calculation. The reason they still matter is practical. Courts have already shown little patience for lawyers who file unverified AI output when the defect could have been found by ordinary checking. A fabricated interest period or an invented statutory rate is not a fake citation, but it creates the same ugly courtroom question: who verified this before filing?
What procurement teams should ask for instead
Astra’s proof results may justify more attention to OpenAI’s technical direction. They do not justify substituting a theorem-proving score for a legal-task reliability record. A litigation or procurement team evaluating AI for legal arithmetic should ask for evidence that fits the task being purchased.
- Independent tests on legal fact patterns, not only vendor demos or general math benchmarks.
- Separate measurements for damages, interest, deadlines, and fee calculations, because those tasks fail in different ways.
- Disclosure of model version, test date, tool access, retrieval settings, calculator or spreadsheet integration, and whether outputs were regenerated after errors.
- Multi-step arithmetic reporting that distinguishes wrong inputs, wrong rule selection, wrong intermediate calculations, and wrong final totals.
- A verification workflow showing who checks the number, what source records are used, how discrepancies are logged, and when a lawyer must review before filing.
- Error examples preserved in audit logs, not only aggregate accuracy claims.
The vendor-claim problem is not unique to legal arithmetic. It is the same procurement pattern that appears whenever a technical capability is marketed into a regulated workflow: the buyer has to separate a system’s advertised capability from independently measured task performance. That distinction also drives the site’s coverage of AI money-laundering detection claims and broader verification workflows. In legal arithmetic, the independent evidence should be boring by design: test sets, version numbers, audit trails, reviewer initials, and reproducible calculations.
The disciplined judgment is straightforward. Astra is a proof-of-capability signal. It is not a filing-safety signal. Until OpenAI or legal-tech vendors produce independent legal-task benchmarks and verification records for real damages, interest, deadline, and fee patterns, AI-computed legal numbers should be treated as supervised nonlawyer output subject to the verification duties lawyers already carry.
References
- Ten advances in mathematics and theoretical computer science, OpenAI, August 1, 2026
- OpenAI announces its next major model Astra by dropping ten previously unsolved math solutions, The Decoder
- OpenAI’s new reasoning AI models hallucinate more, TechCrunch, April 18, 2025
- AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries, Stanford HAI
- Why AI Gets Math Wrong and How to Actually Fix It, Dojo Labs
- There is no free benchmark, PNAS, July 2026
- ABA issues first ethics guidance on a lawyer’s use of AI tools, American Bar Association, July 2024
- Which OpenAI Model for Legal Work?, Debevoise Data Blog, June 22, 2025
- The AI Sanction Wave: $145K in Q1 Penalties Signals Courts Have Lost Patience with GenAI Filing Failures, EDRM, April 2026
Related records
Tool profile
How Meta's AI Spending Reshapes Law Firm ProfitabilityGoverning regulation
Browse the obligations tracker →Preventive workflow
Browse verification workflows →
Report a correction or tip
Spotted an outdated figure, a misstated fact, or a ruling this case record should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.
Report a correction or tip for this record →