Risk data, not endorsement
Evaluations
Citation-accuracy and hallucination-rate benchmarks for named AI legal tools, framed explicitly as risk data rather than product endorsement. Each evaluation discloses its benchmark source, methodology, and test date, and aggregates independent studies rather than vendor-supplied figures where possible. Every tool profile cross-links to the specific Risk Digest cases in which that tool was implicated, turning benchmark scores into traceable risk signals. Serves the procurement and comparison task: is this specific tool safe enough to use. Excludes narrative case reporting (Risk Digest) and procedural steps (Workflows); comparisons must always disclose methodology to avoid misleading side-by-side figures across incompatible test conditions.
Artificial Analysis
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
AuraMate BaziQA(厂商自报)、观象八字实测、DeepOracle对照测试
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
DeepOracle 2026 tool comparison (vendor-run)
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
District of Delaware summary-judgment ruling
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Harvey LAB-AA; Vals AI HLAB
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
PalmReading.pro at Palmist privacy policies; Creen at GPTPalm feature pages
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Princeton LePhantomCite (arXiv preprint)
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Reuters, TechCrunch, CNBC, Legal.io
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
SEC filings, court docket, and regulatory reports
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Stanford Law School; Stanford RegLab; Vals AI; Descrybe
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Stanford/Yale study, arXiv 2024
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
StatusGator; TechTimes
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Suprmind — AI Hallucination Rates & Benchmarks 2026
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Travel China Guide at ChineseNewYear.net
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Vals AI Legal AI Report (VLAIR); Stanford RegLab/HAI
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Vals AI Legal Research Bench; Stanford RegLab/Stanford HAI; DeepMind model card
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Vals AI VLAIR; Stanford RegLab
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Vals Legal Research Bench via BenchLM
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Walang independiyenteng benchmark; self-reported vendor tests lamang (DeepOracle, FateMaster)
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
香港天文台、無相閣
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.




















