Risk data, not endorsement
Evaluations
Citation-accuracy and hallucination-rate benchmarks for named AI legal tools, framed explicitly as risk data rather than product endorsement. Each evaluation discloses its benchmark source, methodology, and test date, and aggregates independent studies rather than vendor-supplied figures where possible. Every tool profile cross-links to the specific Risk Digest cases in which that tool was implicated, turning benchmark scores into traceable risk signals. Serves the procurement and comparison task: is this specific tool safe enough to use. Excludes narrative case reporting (Risk Digest) and procedural steps (Workflows); comparisons must always disclose methodology to avoid misleading side-by-side figures across incompatible test conditions.
AI Business Weekly; Maryland State Bar Association
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Anthropic system card via Mashable; Stanford RegLab
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Artificial Analysis (reported via SCMP)
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Artificial Analysis; Solvimon; Stanford RegLab / Ho et al.
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
AWS Bedrock model-lifecycle table
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
DeepSeek provider tables; BenchLM; Artificial Analysis; CAISI/NIST
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Finance Research Letters (Saggu, Ante, Kopiec 2025)
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Meta Q2 2026 earnings report; Reuters
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
MiniMax API Docs
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
No independent audit located; figures from SITA 2026 insight blog and AeroLogic case study
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
No independent benchmark disclosed; public record review
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
No independent benchmark in H3 public release record
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
No Latency
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Onyx Security product materials and legal sanction records
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
SAOT validation benchmark-gaps analysis
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
SEC filings, company investor relations, Fortune, Reuters, CFO Brew
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Stanford HAI, AI on Trial
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Stanford RegLab/JELS; Stanford HAI; Vals AI LegalBench; Harvey BigLaw Bench; LawNext/VLAIR
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Vals AI VLAIR report
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.
Wukong123
Figures from this source are not directly comparable to other benchmark sources without checking each study's methodology.




















