Risk data, not endorsement
Evaluations
Citation-accuracy and hallucination-rate benchmarks for named AI legal tools, framed explicitly as risk data rather than product endorsement. Each evaluation discloses its benchmark source, methodology, and test date, and aggregates independent studies rather than vendor-supplied figures where possible. Every tool profile cross-links to the specific Risk Digest cases in which that tool was implicated, turning benchmark scores into traceable risk signals. Serves the procurement and comparison task: is this specific tool safe enough to use. Excludes narrative case reporting (Risk Digest) and procedural steps (Workflows); comparisons must always disclose methodology to avoid misleading side-by-side figures across incompatible test conditions.
Artificial Lawyer
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Artificial Lawyer (Legal Complex)
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Axios
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Bloomberg Law
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Cambridge International Journal of Legal Information
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
CJEU Superleague judgment, European Commission warning, Swiss Verein analysis, House Judiciary investigation
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
CoreWeave Q1 2026 earnings report
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Google DeepMind model card and launch materials (self-reported)
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
HAQQ, Percipient, Legal Benchmarks, Snorkel AI, Suprmind
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
LegalOn, Vals AI, HAQQ, Stanford RegLab, Harvey, Clio, Suprmind/AA-Omniscience
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
LePhantomCite; Vals AI Legal Research Bench 2026
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
NIST/CAISI; Artificial Analysis; Vectara; BenchLM
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Source undisclosed
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
UpdatedHallucination rateNot measured / undisclosedWhy DeepSeek V4 Flash API pricing is the cheap part
Compare DeepSeek V4 Flash token rates against what legal AI platforms actually charge per seat, and see why raw model cost is the smallest line in a legal AI budget — verified legal data, security controls, and professional-risk protections carry the real expense.
UpdatedHallucination rateNot measured / undisclosedGPT-5.6 Leads DeepSeek V4 on Legal Research, Neither Is Safe
The July 2026 legal-research benchmarks show GPT-5.6 leading DeepSeek V4 by a wide margin, yet the best generalist still fails more than half of strict checklist-style tasks. This evaluation turns the scores into a procurement and filing risk judgment, mapping each model family's distinct failure modes and the verification obligations that survive either choice.
Stanford HAI
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Stanford RegLab and HAI
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Stanford RegLab/HAI; Vals VLAIR
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Stanford RegLab/Stanford HAI; HAQQ; Vals AI VLAIR
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Supreme Court dockets, SCOTUSblog, AP, NPR
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
xAI launch benchmarks
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.


















