Benchmarks

Citation-accuracy and reliability benchmarks for AI legal research and drafting tools, framed as risk assessment rather than product endorsement. Each entry discloses the benchmark source type (peer-reviewed academic study, vendor-published report, or independent test), the tool name and version tested, the test date, and a methodology summary, and visually separates independent findings from vendor-coordinated ones. Entries cross-link to any Risk Digest case naming the same tool. Does not contain case outcomes (Risk Digest) or how-to guidance (Workflows), though it links to both.

ToolVendorSourceHallucination rateVersion testedTest date
Mugshot ruling tested in French mass shooting case
California NCAA Age Lawsuit Tests Multi-Jurisdiction Strategy
Understanding New York's 2026 car accident injury law changes
Ortega's 2025 No-Election Rule: Legal Implications for the Profession
How Nintendo's Palworld Payout Caps at $30,000
The Defamation Calculus Behind Panettiere's Anonymous Allegations
Family sues after Park Police policy change leads to fatal chase
Purpose-Built Legal AI Outperforms ChatGPT for Contract Review
The SAVE America Act Mandates Unregulated AI in Election Administration
Spellbook AI Contract Drafting Tool: Evaluation for Legal TeamsRally Legal (Spellbook)
Why Wildfire Containment Percentages Matter in Court
Legal implications of the AA2653 engine failure for airlines
What AI Paralegal Tools Still Get Wrong: Six Failure Modes Beyond Hallucinations
What Legally Protects Journalists From Subpoenas Like the Air Force One Case?
Who Bears Legal Liability After Alaska Airlines Mechanical Failure?
Andrew and Tristan Tate criminal case overview explained
The pleading strategy behind Apple's lawsuit against OpenAI
Best Legal AI Tools by Practice Need: A Use-Case-Based Comparison Guide
Bill Pulte's Acting DNI Appointment Tests the Vacancies Reform Act
How Texas Law Applies to the Charles Medina Road Rage Death
Blogarama - Blog Directory