Skip to content

Risk data, not endorsement

Evaluations

Citation-accuracy and hallucination-rate benchmarks for named AI legal tools, framed explicitly as risk data rather than product endorsement. Each evaluation discloses its benchmark source, methodology, and test date, and aggregates independent studies rather than vendor-supplied figures where possible. Every tool profile cross-links to the specific Risk Digest cases in which that tool was implicated, turning benchmark scores into traceable risk signals. Serves the procurement and comparison task: is this specific tool safe enough to use. Excludes narrative case reporting (Risk Digest) and procedural steps (Workflows); comparisons must always disclose methodology to avoid misleading side-by-side figures across incompatible test conditions.

Artificial Lawyer

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Artificial Lawyer (Legal Complex)

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Axios

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Bloomberg Law

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Cambridge International Journal of Legal Information

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

CJEU Superleague judgment, European Commission warning, Swiss Verein analysis, House Judiciary investigation

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

CoreWeave Q1 2026 earnings report

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Google DeepMind model card and launch materials (self-reported)

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

HAQQ, Percipient, Legal Benchmarks, Snorkel AI, Suprmind

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

LegalOn, Vals AI, HAQQ, Stanford RegLab, Harvey, Clio, Suprmind/AA-Omniscience

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

LePhantomCite; Vals AI Legal Research Bench 2026

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

NIST/CAISI; Artificial Analysis; Vectara; BenchLM

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Source undisclosed

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Stanford HAI

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Stanford RegLab and HAI

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Stanford RegLab/HAI; Vals VLAIR

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Stanford RegLab/Stanford HAI; HAQQ; Vals AI VLAIR

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Supreme Court dockets, SCOTUSblog, AP, NPR

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

xAI launch benchmarks

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.