Skip to content

Evaluations

Neither Codex nor Claude Code Is Safe for Legal AI Agents

Legal-tech buyers weighing OpenAI Codex against Anthropic Claude Code get a benchmark-anchored risk comparison: the 2026 Princeton LePhantomCite study shows the two platforms fail legal citation verification in opposite ways, and neither reaches a delegation-safe threshold. The verdict is a mandatory human verification gate under either choice, not a recommendation to standardize on one tool.

By Editorial TeamPublished Aug 26, 2026
Tool
OpenAI Codex, Anthropic Claude Code
Benchmark source
Princeton LePhantomCite (arXiv preprint)
Hallucination rate
LePhantomCite recall: Claude Code 62.8%; GPT-5 agentic 84.4%
Test methodology
Citation-hallucination detection benchmark across CourtListener absence, incorrect pincites, and content misrepresentation
Test date
Aug 6, 2026

For buyers comparing Codex and Claude Code for legal AI agents, the 2026 record does not support a clean platform winner. The strongest direct evidence, Princeton’s LePhantomCite citation-verification benchmark, shows two different ways to fail a legal review workflow: Claude Code with Opus 4.8 is much cleaner when it flags a hallucinated citation, while the best GPT-5 agentic verification system catches more hallucinations but generates far more noise. Neither result is delegation-safe for citation-bearing legal work.

Two AI systems compare highlighted footnotes across a law-office desk

The direct evidence: LePhantomCite

LePhantomCite matters because it is not a general legal-reasoning leaderboard being stretched into a procurement claim. It is a direct citation-hallucination detection test, published as an arXiv preprint submitted on June 19, 2026 and revised on August 6, 2026. That status matters: it is not peer-reviewed, and the reported platform results should be treated as controlled benchmark evidence rather than a settled ranking. Still, it is the load-bearing comparison available in 2026 for Claude Code against an OpenAI GPT-5 agentic verification setup on a legal citation task.[1]

LePhantomCite headline results for hallucinated-citation detection.[1]
System testedF1PrecisionRecallPractical reading
Claude Code + Opus 4.868.8%76.1%62.8%Fewer bad flags; misses more hallucinated citations overall
Best GPT-5 agentic verification system55.0%40.8%84.4%Catches more hallucinations; sends many more false alarms to reviewers

The most useful number in that table is not necessarily F1. A law firm does not file an F1 score. It files a brief after someone has decided whether the tool’s flags are reliable enough to govern human attention.

LePhantomCite’s false-positive contrast makes the workflow problem clearer. On citations absent from CourtListener, Claude Code produced a 10.9% false-positive rate, compared with 25.0% for GPT-5 and 65.9% for Gemini 2.5 Flash. The paper attributes Claude Code’s lower false-positive behavior to treating database absence as a coverage gap rather than proof that a citation is fabricated.[1]

That is the difference between a verification assistant and a panic machine. A missing CourtListener hit can mean the citation is wrong. It can also mean the database does not have the case, the reporter form differs, the citation is malformed but recoverable, or the relevant authority sits outside the tested coverage. An agent that collapses those possibilities into “fake” may look vigilant in a demo and become expensive in a filing room.

Recall is not free, and precision is not safety

A legal document with many noisy alert markers beside a document with fewer precise markers

The GPT-5 agentic result is attractive for the one thing no supervising lawyer wants to miss: fabricated legal support. Its 84.4% recall means it found a larger share of hallucinated citations in the benchmark than Claude Code did. In a vacuum, higher recall sounds like the safer posture.[1]

But legal review is not a vacuum. Every false positive becomes a task for someone who must open the source, check the reporter, check the pincite, check the quoted proposition, and decide whether the tool is wrong or the draft is wrong. At 40.8% precision, the best GPT-5 agentic system is asking reviewers to sort through substantially more non-errors than Claude Code does at 76.1% precision.[1]

That burden is not just an efficiency problem. Noise changes behavior. If an associate sees enough bad flags, the verification process starts to become selective: the easy flags get checked, the annoying flags get postponed, and the “probably fine” flags become a second-order hallucination risk. The tool has not merely failed to save time; it has taught the reviewer to distrust the channel that is supposed to surface danger.

Claude Code’s risk runs in the opposite direction. Its precision is markedly better, which means a flag from Claude Code is more likely to deserve attention. But its 62.8% recall means it missed more hallucinated citations overall than the GPT-5 agentic system. A cleaner queue is not the same thing as a clean brief.[1]

So the operational choice is not “Codex for safety” versus “Claude Code for safety.” It is whether the firm wants to manage a high-noise, higher-recall queue or a lower-noise, lower-recall queue. Both require a human gate before filing.

The subcategory failures are the harder procurement problem

The headline precision-recall split is only the first pass. The more uncomfortable part of LePhantomCite is what happens when the errors become legally subtler: incorrect pincites and content misrepresentation.

Selected LePhantomCite subcategory recall results.[1]
Error typeClaude Code recallGPT-5 recallWhy it matters
Incorrect pincites3.6%52.8%A cited case may exist, but the pinpoint does not support the proposition
Content misrepresentation66.4%83.2%The cited source exists, but the brief distorts what it says

The incorrect-pincite result should be handled carefully. LePhantomCite reports a wide confidence interval for Claude Code’s incorrect-pincite recall, 3.6% ±14.0, so the number is better read as a fragility warning than as a stable estimate of real-world performance. It is still the wrong shape for delegation: the very error that often survives superficial citation checks is not reliably caught.[1]

Content misrepresentation is worse because the paper identifies it as the hallucination type most likely to distort legal outcomes. This is the error where a cited authority exists, the citation may look plausible, and the proposition attributed to the authority is still wrong. A reviewer cannot clear that by confirming that the case is real. Someone has to read the relevant passage and compare it to the sentence in the brief.[1]

GPT-5’s higher recall on content misrepresentation is meaningful. It is also not enough. Missing roughly one in six of that category on the benchmark is not a filing-safe threshold, and the system’s broader false-positive profile means the missed errors would sit inside a review environment already fighting alert fatigue.[1]

Claude Code’s cleaner flagging behavior does not solve this either. A reviewer who relies on Claude Code to identify the dangerous places may receive fewer distractions, but the benchmark still leaves too many hallucinations outside the flagged set. That is especially serious for citation-bearing drafts, where a single unflagged false proposition can be more consequential than a stack of harmless formatting errors.

The sanction examples are reality checks, not anecdotes to wave away

The reason this benchmark matters is not that legal academics have invented an artificial task. Courts and judges have already seen the failure class. In 2025, TechCrunch reported that Anthropic’s own outside counsel missed a Claude-hallucinated citation in an expert-report context, with Judge Susan van Keulen ordering a response.[2]

In 2026, Reuters reported that Sullivan & Cromwell apologized to a federal judge over AI-generated inaccurate citations in a filing. The available research basis supports that general description, but not fine-grained claims about the internal review sequence, so the point should stay narrow: sophisticated counsel can still let citation errors pass when the verification gate fails.[3]

And Mata v. Avianca remains the foundational warning because it turned fake case citations from a legal-tech embarrassment into a sanctions event. The lesson is not that every hallucinated citation will produce sanctions. It is that “the tool said it was real” has already proved useless as a professional-responsibility answer when a filing reaches the court.[4]

What Vals and Endor add, and what they do not

Vals AI’s Legal Research Bench points in a direction that is broadly consistent with caution, but it is not the same kind of evidence as LePhantomCite. Vals is model-level legal-research evidence, not a direct Codex-versus-Claude Code citation-verification procurement test. On its strict all-pass measure, Claude Opus 5 leads GPT-5.6 Sol, 55.29% to 48.08%, and Vals identifies reconciliation of conflicting authority as the dominant failure mode.[5]

That matters because legal research failures often do not look like obviously fake citations. They look like overreading favorable authority, underweighting contrary authority, or failing to reconcile tension across cases. A citation verifier that confirms existence does not cure a research agent that has misunderstood the relationship among authorities.

The platform mapping still has to be labeled for what it is: synthesis. Claude Opus 5 performance on a model benchmark is relevant to evaluating Claude Code, and GPT-5.6 Sol performance is relevant to evaluating an OpenAI-side agentic stack, but those scores do not prove how either product will behave in a firm’s document-management system, citation database environment, instruction layer, or review protocol.

Endor Labs’ Agent Security League is even farther from citation reliability, but it becomes relevant when legal teams use coding agents to build internal tools, intake automations, filing-support scripts, or research workflow glue. In the July 31, 2026 update, Claude Code with Opus 5 scored 32.4% secure, while Codex with GPT 5.6 Sol scored 23.5% secure. Even the leader left roughly two-thirds of agent-produced code failing the security tests.[6]

That is not proof that Claude Code is more reliable for legal citation checking, and it should not be sold that way. It is a separate risk signal for teams treating legal AI agents as builders of internal legal automation. A tool that drafts code insecurely and verifies citations imperfectly needs controls in both places, not a single procurement slogan.

A defensible permission rule for citation-bearing drafts

A human hand stamps verified on legal documents after an AI scanning step

A law firm can evaluate either platform as an assistive layer. It can use a GPT-5 agentic system to widen the net for possible hallucinations. It can use Claude Code where a lower-noise flagging channel is more valuable. It can run both in a comparative pilot and measure reviewer burden, missed-error rate, false-positive rate, source-coverage gaps, and how often reviewers override the tool.

What the 2026 evidence does not support is authorizing either platform to clear citation-bearing legal work without human verification. The permission rule should be blunt: an AI agent may identify citations for review, prioritize suspicious passages, and produce a verification memo; it may not be the final authority that a citation exists, that the pincite supports the proposition, or that the quoted authority has not been misrepresented.

For teams already tracking OpenAI-side legal reliability, the adjacent ChatGPT legal reliability record and the Grok 4.6 versus GPT-5.6 Sol benchmark comparison are useful background. For the verification gate itself, the practical next document is the six-phase AI hallucination audit checklist. The procurement decision should tie tool permission to that kind of audit trail, not to a benchmark delta that disappears before the filing tray.

References

  1. Who Checks the Citations? Benchmarking Legal Hallucination Detection, arXiv, revised August 6, 2026.
  2. Anthropic’s lawyer apologizes for using Claude to file a hallucinated citation, TechCrunch, May 15, 2025.
  3. Sullivan & Cromwell apologizes to federal judge over AI-generated inaccurate citations, Reuters, April 21, 2026.
  4. Mata v. Avianca, Inc. Sanctions Opinion, U.S. District Court for the Southern District of New York, June 22, 2023.
  5. Legal Research Bench, Vals AI.
  6. AI Code Security Benchmark, Endor Labs, updated July 31, 2026.

Report a correction or tip

Spotted an outdated figure, a misstated fact, or a ruling this tool profile should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.

Report a correction or tip for this record →
Blogarama - Blog Directory