Skip to content
Lex Machina Review logoLex Machina Review
Menu

Risk Digest

What the '1933 double' Reveals About ChatGPT Benchmarks

SCOTUSblog's 2023 and 2025 ChatGPT tests improved from 21/50 to as high as 45/50, yet the retest produced new fabricated citations — benchmark gains are not verified legal output. The 'Justice James F. West impeached in 1933' error anchors the analysis, and the real '1933 double' coin litigation (Langbord) is a separate, non-AI matter.

REPORTED — UNVERIFIED
Jurisdiction
US federal
Court
U.S. Supreme Court
AI tool named
ChatGPT
Ruling date
Mar 21, 2025
Source document
View primary court order ↗
Last verified
Aug 4, 2026

Lex Machina Review is an independent risk-tracking and reference resource. Nothing on this site is legal advice, and using it does not create an attorney-client relationship. Every record is reviewed against primary sources but may not reflect the most current status of a matter — always verify directly against the cited court order, rule text, or a licensed attorney before relying on it.

Companion explanation — secondary to the source document above

The useful part of SCOTUSblog’s two Supreme Court quizzes is not that the score went up. It is that the wrongness changed shape. In 2023, ChatGPT got 21 of 50 right and still produced the bizarre claim that “Justice James F. West” was impeached in 1933; in 2025, the scores rose to 29/50 on 4o, 36/50 on o3-mini, and 45/50 on o1, but the retest replaced that old West/1933 error with new fabrications, including an invented Chief Justice Roberts quote and an invented Janus relisting history. [1][2]

SCOTUSblog January 2023 benchmark article header image

What the 2023 test actually showed

The 2023 post is still the clearest diagnostic because its error is so cheap to check. There was no Justice James F. West. There was no Supreme Court impeachment in 1933. The model appears to have blended a real constitutional-history slot with invented identity details and a false year, which is exactly the kind of answer that can look plausible until someone checks a short official list. SCOTUSblog reported 21 correct answers, 26 incorrect answers, and 3 incomplete answers. [1]

The phrase 1933 double is noisy in search results because it often points to the 1933 Double Eagle dispute, which is a real federal case record but a separate matter from this AI benchmark. It should not be confused with the fabricated 1933 impeachment claim.

SCOTUSblog March 2025 benchmark retest header image

What changed in 2025

The retest is where the procurement question gets interesting. ChatGPT 4o scored 29/50, o3-mini scored 36/50, and o1 scored 45/50. The old West/1933 fabrication was gone, which matters, but the 2025 run still invented other legal details instead of converging on something you would call verified output. SCOTUSblog identified an invented Roberts quote and an invented Janus v. AFSCME relisting history in that retest. [2]

That is a real improvement in the narrow sense that fewer answers are plainly wrong. It is not a certification that the model can be trusted to preserve a legal fact from an answer to a filing.

Why the 1933 claim is such a clean warning

A year-plus-person claim about court history should trigger a fast source check because the record is finite and official. That is why the West/1933 error is more useful as a benchmark failure than as a joke: once the model names a judge, a year, and an event that either happened or did not happen, the mismatch is cheap to verify. The fact that the 2025 test dropped that exact fabrication does not tell you the class of error is gone; it only tells you the model learned not to repeat one memorable mistake. [1][2]

Why filing risk still belongs in the conversation

That distinction matters because the real consequence case is not hypothetical. In Mata v. Avianca, Judge P. Kevin Castel sanctioned Steven Schwartz, Peter LoDuca, and Levidow, Levidow & Oberman a total of $5,000 on June 22, 2023 after six fictitious ChatGPT-generated citations and what the court described as “acts of conscious avoidance and false and misleading statements to the court.” [3]

Mata is not the same thing as a benchmark quiz, but it is the bridge between trivia and consequences. One tells you the tool can answer more Supreme Court questions correctly than before; the other shows that fabricated legal material can still make it into a filing if nobody stops it.

The broader scale signal is also worth keeping in view. Damien Charlotin’s AI Hallucination Cases Database lists 1,811 decisions worldwide, including 1,252 in the United States, as of its July 29, 2026 update. That is enough to show the problem is recurring, but the page changes often enough that the count should be rechecked at publication rather than copied as if it were static. [4]

So the clean reading is modest. Rising benchmark scores reduce some visible errors and make the tool less random. They do not turn a benchmark into verified legal output, and they do not retire the old rule that anachronistic year-plus-person claims about court history deserve a source check before anyone relies on them in a memo or brief.

References

  1. “No, Ruth Bader Ginsburg did not dissent in Obergefell — and other things ChatGPT gets wrong about the Supreme Court.” SCOTUSblog. Jan. 26, 2023. Link
  2. “We’re not there to provide entertainment...” — ChatGPT and the Supreme Court, two years later.” SCOTUSblog. Mar. 21, 2025. Link
  3. “New York lawyers sanctioned for using fake ChatGPT cases in legal brief.” Reuters. June 22, 2023. Link
  4. “AI Hallucination Cases Database.” Damien Charlotin. Last updated July 29, 2026. Link

Report a correction or tip

Spotted an outdated figure, a misstated fact, or a ruling this case record should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.

Report a correction or tip for this record →
Blogarama - Blog Directory