The Ethics Landscape: ABA Opinions, State Bar Guidance, and Court Orders on AI Legal Research
Before comparing individual products, it is essential to understand the professional responsibility framework that governs any use of AI in legal practice. The ABA Model Rules — particularly Rule 1.1 (competence), Rule 1.6 (confidentiality), and Rule 5.3 (supervision of nonlawyer assistance) — have been repeatedly interpreted to require that lawyers understand the capabilities and limitations of the technology they employ. A growing number of state bar ethics opinions have directly addressed generative AI, and federal judges have begun issuing standing orders that mandate disclosure and verification of AI-generated content in court filings.
The duty of technology competence under Rule 1.1, clarified in ABA Formal Opinion 512 (2024), now explicitly requires lawyers to stay informed about the benefits and risks of AI tools. When a research platform returns a fabricated citation or a misstatement of law, the lawyer — not the software — bears the consequences. This makes the verification infrastructure built into each platform not a convenience feature but a core ethical safeguard.
What the Stanford RegLab Study Actually Found (and What Changed Since 2024)
The most frequently cited accuracy benchmark for legal AI research tools is the May 2024 study from Stanford's RegLab, which tested over 200 open-ended legal queries across multiple platforms. The study found that Lexis+ AI hallucinated approximately 17% of the time, Westlaw AI-Assisted Research hallucinated approximately 34%, and Ask Practical Law AI also hallucinated 17%. For context, general-purpose chatbots (ChatGPT, Claude, Gemini) hallucinated between 58% and 82% on the same legal queries.
However, a critical caveat applies: the Stanford study tested products that no longer exist under those names. Lexis+ AI was rebranded to Lexis+ with Protégé in February 2026, and Westlaw AI-Assisted Research has been superseded by CoCounsel Deep Research, launched in August 2025. Neither 2026 product has been subjected to an independent, peer-reviewed benchmark. The 17% and 34% figures are directional, not current.
| Product Tested (May 2024) | Hallucination Rate | Current 2026 Equivalent |
|---|---|---|
| Lexis+ AI | 17% | Lexis+ with Protégé |
| Westlaw AI-Assisted Research | 34% | CoCounsel Deep Research |
| Ask Practical Law AI | 17% | Part of CoCounsel suite |
| General-purpose chatbots (average) | 58–82% | N/A |
How Legal RAG Systems Hallucinate: Retrieval Failure, Inapplicable Authority, and Sycophancy
The Stanford RegLab study categorized hallucinations into two types — incorrect law and misgrounded citations — and identified three underlying failure modes in retrieval-augmented generation (RAG) systems used by legal AI tools. Understanding these failure modes is essential for designing effective verification workflows.
- Retrieval failure: The system fails to locate the relevant legal authority in its database, so the language model generates an answer from its pre-training data, which may be outdated or entirely fabricated.
- Inapplicable authority: The system retrieves a real case or statute, but the authority is no longer good law or does not apply to the jurisdiction or factual scenario. The study gave the example of a system citing the "undue burden" standard from abortion precedent that had been overruled post-Dobbs.
- Sycophancy: The model agrees with a user's false or misleading premise. The study documented a case where a user asked whether Justice Ginsburg dissented in Obergefell v. Hodges, and the model fabricated a dissent — likely because it was predisposed to agree with the user's implied claim.

CoCounsel's Verification Infrastructure: KeyCite Flags, Hallucination Checker, and Litigation Document Analyzer
Thomson Reuters CoCounsel, built on Westlaw's editorial backbone, provides multiple layers of verification that directly support a lawyer's duty of competence. The foundation is Westlaw's KeyCite system, which has been maintained by more than 1,200 full-time attorney-editors for over 100 years. KeyCite assigns color-coded flags to every cited authority:
- Red flag — the case is no longer good law (e.g., reversed, overruled).
- Yellow flag — the case has negative history but has not been directly reversed.
- Red-striped flag — the case has been partially overruled on some points.
- Blue-striped flag — the case is involved in federal appeals proceedings.
- Overruling Risk icon — a warning that a key precedent may be at risk of being overturned.
Beyond KeyCite, CoCounsel's Litigation Document Analyzer (LDA) now includes a dedicated hallucination checker designed specifically to identify citations that may not be real. The LDA also performs language analysis that compares assertions in a brief against the cited source material to surface potential mischaracterizations. CoCounsel Deep Research, described by Thomson Reuters as the legal industry's first professional-grade agentic AI research capability, creates iterative research plans and delivers reports with transparent reasoning, grounding every output exclusively in Westlaw and Practical Law content.
At the time of writing, over 20,000 law firms and corporate legal departments use CoCounsel, including a majority of Am Law 100 firms, according to Thomson Reuters' August 2025 announcement.
Lexis+ with Protégé's Verification Infrastructure: Shepard's, Legal AI vs General AI Modes, and the Vault
Lexis+ with Protégé (formerly Lexis+ AI, rebranded February 2026) takes a dual-mode approach that gives users explicit control over the grounding of AI output. The platform offers two distinct modes:
- Legal AI mode — optimized for legal research and drafting, grounded exclusively in LexisNexis's editorial corpus. Every citation includes Shepard's signal indicators (red, yellow, green, blue) that show whether a case is still good law. Shepard's functions similarly to KeyCite, with the addition of detailed negative treatment analysis and citation history.
- General AI mode — provides access to large language models from OpenAI, Google, and Anthropic for broader reasoning and exploration. This mode is not grounded in legal sources and carries a much higher risk of hallucination. The mode is clearly labeled, but the burden falls on the user to verify all output.
Lexis+ with Protégé also includes the Vault, a secure document storage system with configurable retention policies. Documents uploaded individually (up to 10) are purged at session end; bulk uploads of 11 or more documents create a Vault that persists until deleted. Critically, LexisNexis states that customer data is not used to train public AI models, addressing confidentiality concerns under Model Rule 1.6.
Real-World Consequences: Over 1,300 Sanctions Cases Globally for Unverified AI-Generated Filings
The risk of unverified AI output is not theoretical. According to a Thomson Reuters redline comparison blog citing the work of Damien Charlotin, over 1,300 identified cases worldwide document sanctions for unverified AI-generated court filings. Penalties range from fines of thousands to six figures, dismissals, public reprimands, filing restrictions, and disciplinary referrals to state bars.
High-profile cases such as Mata v. Avianca (2023), where counsel submitted a brief containing fabricated citations generated by ChatGPT, have become cautionary tales taught in law schools. But the risk extends beyond generative AI to AI-assisted legal research tools — if a platform returns a misgrounded citation and the attorney files it without verification, the professional liability falls on the attorney.
| Sanction Type | Example Range | Frequency |
|---|---|---|
| Fines | $5,000 – $100,000+ | Common |
| Dismissal or default judgment | Case ending | Documented in multiple federal courts |
| Public reprimand | Permanent record | Increasing |
| Disciplinary referral | State bar investigation | Growing trend |
Practical Verification Workflows: How to Audit AI Output on Both Platforms
Regardless of which platform you use, the verification workflow must include the same core steps. The differences between CoCounsel and Lexis+ with Protégé lie in the specific tools available to execute each step.
- Check every cited authority using KeyCite (CoCounsel) or Shepard's (Lexis+ with Protégé). Look for red, yellow, or striped flags. A green signal does not guarantee the citation supports the proposition — it only means the case has not been negatively treated.
- Open the source document and read the specific passage the AI claims supports its statement. Do not rely on the AI's summary. Confirm the holding, reasoning, and jurisdiction align with your argument.
- Use CoCounsel's hallucination checker (in Litigation Document Analyzer) to flag citations that may not be real. On Lexis+ with Protégé, stay in Legal AI mode and verify that every Shepard's link resolves to a real document.
- Cross-reference secondary sources such as American Law Reports, treatises, or practice guides available in each platform. CoCounsel's Deep Research includes Practical Law annotations; Lexis+ with Protégé offers LexisNexis practice guides.
- Document your verification process. If a court asks about your diligence, you should be able to demonstrate that you checked each source, reviewed the original text, and confirmed the proposition.
Scorecard: Which Platform Reduces Ethical Risk More Effectively?
The following table summarizes the key differences between CoCounsel and Lexis+ with Protégé as they relate to professional responsibility obligations. Use it as a reference when evaluating which platform aligns with your firm's risk tolerance and practice needs.

| Evaluation Criterion | CoCounsel (Thomson Reuters) | Lexis+ with Protégé (LexisNexis) |
|---|---|---|
| Independent hallucination data (2026) | No recent peer-reviewed study | No recent peer-reviewed study |
| Legacy benchmark (May 2024) | 34% for predecessor (Westlaw AI-Assisted Research) | 17% for predecessor (Lexis+ AI) |
| Citation verification system | KeyCite with 5 flag types plus Overruling Risk | Shepard's with signal indicators and detailed negative treatment |
| Dedicated hallucination detection | Litigation Document Analyzer hallucination checker | Not explicitly documented; relies on Legal AI mode grounding |
| Deep research accuracy safeguards | Agentic Deep Research grounded in Westlaw only | Deep research available only with web sources selected in Best Fit mode (per Thomson Reuters comparison) |
| User-adjustable confidence modes | Single mode grounded in Westlaw/Practical Law | Dual mode: Legal AI (grounded) and General AI (ungrounded) |
| Data confidentiality posture | Not used to train public models (stated) | Customer data not used to train public models (stated) |
| Document retention control | Standard Westlaw session storage | Vault with configurable retention policies |
For practitioners in high-stakes litigation or jurisdictions with strict verification requirements, CoCounsel's integrated hallucination checker and longstanding KeyCite editorial layer provide a strong safety net. The platform's single-mode design removes the temptation to use an ungrounded AI mode for legal research.
For firms that value flexibility — for example, using AI for brainstorming or general legal reasoning alongside research — Lexis+ with Protégé's dual-mode setup offers clear utility, provided the lawyer stays in Legal AI mode for research tasks. The risk lies in mode confusion: an attorney who absentmindedly uses General AI mode for a research query may receive output that appears plausible but is not grounded in any legal source.
The ultimate scorecard, however, is not a feature list — it is the verification workflow the lawyer actually follows. Both platforms can be used ethically if every AI-generated citation is verified against the original source. And both platforms can be used unethically if output is filed without review. The technology reduces the burden of verification, but it does not eliminate the lawyer's duty to ensure that the law presented to a court is accurate.