Skip to content

Risk data, not endorsement

Evaluations

Citation-accuracy and hallucination-rate benchmarks for named AI legal tools, framed explicitly as risk data rather than product endorsement. Each evaluation discloses its benchmark source, methodology, and test date, and aggregates independent studies rather than vendor-supplied figures where possible. Every tool profile cross-links to the specific Risk Digest cases in which that tool was implicated, turning benchmark scores into traceable risk signals. Serves the procurement and comparison task: is this specific tool safe enough to use. Excludes narrative case reporting (Risk Digest) and procedural steps (Workflows); comparisons must always disclose methodology to avoid misleading side-by-side figures across incompatible test conditions.

Artificial Analysis

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

AuraMate BaziQA(厂商自报)、观象八字实测、DeepOracle对照测试

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

DeepOracle 2026 tool comparison (vendor-run)

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

District of Delaware summary-judgment ruling

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Harvey LAB-AA; Vals AI HLAB

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

PalmReading.pro at Palmist privacy policies; Creen at GPTPalm feature pages

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Princeton LePhantomCite (arXiv preprint)

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Reuters, TechCrunch, CNBC, Legal.io

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

SEC filings, court docket, and regulatory reports

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Stanford Law School; Stanford RegLab; Vals AI; Descrybe

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Stanford/Yale study, arXiv 2024

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

StatusGator; TechTimes

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Suprmind — AI Hallucination Rates & Benchmarks 2026

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Travel China Guide at ChineseNewYear.net

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Vals AI Legal AI Report (VLAIR); Stanford RegLab/HAI

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Vals AI Legal Research Bench; Stanford RegLab/Stanford HAI; DeepMind model card

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Vals AI VLAIR; Stanford RegLab

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Vals Legal Research Bench via BenchLM

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

Walang independiyenteng benchmark; self-reported vendor tests lamang (DeepOracle, FateMaster)

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.

香港天文台、無相閣

Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.