Is CoCounsel Safe for Law Firm Legal Research?
Independent benchmarks show CoCounsel excels at closed-document analysis, but no independent benchmark covers its agentic research mode, and a documented sanction record shows research outputs cannot be filed unverified. This evaluation provides a risk-signal scorecard law firms can use to decide where CoCounsel is safe to rely on and where verification is mandatory.
- Tool
- CoCounsel
- Benchmark source
- Vals AI Legal AI Report (VLAIR); Stanford RegLab/HAI
- Hallucination rate
- Not measured / undisclosed
- Test methodology
- VLAIR task-specific closed-document benchmarks; Stanford RegLab open-research hallucination tests on pre-relaunch Thomson Reuters products
For a law firm evaluating CoCounsel as a Westlaw AI legal research tool, the useful first question is not whether CoCounsel is “good.” It is which task the firm is asking it to perform, what evidence exists for that task, and who is going to sign the filing if the answer is wrong. This is an editorial tool-reliability evaluation, not legal advice.
| Use case | Current risk signal | What the evidence supports | Required firm posture |
|---|---|---|---|
| Closed-document analysis: summarizing, extracting, answering questions from provided materials | Favorable | CoCounsel’s VLAIR results are strong on bounded document tasks: 79.5% average across its four entered tasks, 77.2% in document summarization, and lawyer-baseline outperformance in Document Q&A, Summarization, and Data Extraction [1]. | Reasonable to use with ordinary professional review, source-bounded queries, and spot checks tied to the uploaded record. |
| Open-ended legal research and agentic Deep Research | Mandatory verification | The available independent benchmark record does not establish reliability for CoCounsel’s current agentic research mode. Earlier Stanford testing measured Westlaw AI-Assisted Research and Ask Practical Law AI, not the later agentic CoCounsel Legal relaunch [3][4]. | Treat every case, quotation, holding, and procedural statement as unverified until checked against primary sources. |
| Security, confidentiality, and integrations | Limited procurement signal | Thomson Reuters states that CoCounsel Legal has SOC 2 and ISO 42001 controls, does not use client data to train underlying models, and integrates with Westlaw, Practical Law, iManage, NetDocuments, SharePoint, HighQ, and Microsoft 365 [10]. | Useful for vendor review, but do not confuse confidentiality controls with research accuracy. |
| Pricing and commercial transparency | Attributed but not confirmed | There is no official public price sheet in the available record. Third-party 2026 estimates describe CoCounsel as a Westlaw add-on with a broad monthly per-user range, but those figures are not Thomson Reuters-confirmed [13][14]. | Require written pricing, seat definitions, data terms, renewal terms, and scope of included Westlaw/Practical Law access before procurement approval. |

The strong part: bounded document work
The best independent evidence for CoCounsel is not a marketing demo. It is the Vals AI Legal AI Report, usually referred to as VLAIR, and the companion CoCounsel application report. In that testing, CoCounsel entered four of the seven available tasks and averaged 79.5% across those entered tasks [1].
The task-level results matter more than the average. CoCounsel scored 89.6% in Document Q&A against a 70.1% lawyer baseline, 77.2% in Summarization against a 50.3% lawyer baseline, and 73.2% in Data Extraction against a 71.1% lawyer baseline. It trailed the lawyer baseline only in Chronology Generation, 78% to 80.2% [1].
Those are not trivial numbers. They describe legal work that firms actually need help with: reading a known record, pulling out information from known materials, and answering questions from a bounded source set. A knowledge-management lawyer can design review around that. The user can preserve the source set, inspect the cited passage, compare the extraction to the document, and correct the output before it leaves the team.
That is why CoCounsel’s VLAIR performance is a favorable signal for closed-document analysis. It does not make the tool self-verifying, and it does not eliminate lawyer review. But it is evidence that the product performed well on tasks where the answer is supposed to come from material the user has supplied.
Do not overread VLAIR
VLAIR is useful precisely because it is task-specific. It becomes less useful when treated as a single league table for “best legal AI.” Vendors did not all enter the same task set. CoCounsel entered four of seven tasks, while LexisNexis withdrew from all non-research tasks, so broad average-to-average comparisons across vendors would be methodologically sloppy [2].
The clean reading is narrower: CoCounsel has strong independent evidence on the closed-document tasks it entered. That evidence maps well onto summarizing a deposition transcript, extracting terms from a contract set, or asking questions of a provided record. It does not, by itself, establish that CoCounsel can safely perform open-ended legal research for a court filing.
The distinction is not academic. Closed-document analysis starts with a defined universe. Legal research asks the system to find, select, characterize, and often synthesize authorities from a much larger legal universe. A wrong answer in the first setting may be caught by comparing the output to the uploaded document. A wrong answer in the second setting can become a fake quotation, a misstated holding, or a non-existent case in a filed brief.
The research-mode evidence gap
The Stanford RegLab and HAI study is important, but it has to be kept in its lane. The study reported that Westlaw AI-Assisted Research hallucinated in more than 34% of responses, while Ask Practical Law AI and Lexis+ AI hallucinated in more than 17% of responses; the work was later peer-reviewed in the Journal of Empirical Legal Studies in 2025 [3][4].
Those figures should not be quoted as CoCounsel’s hallucination rate. The study covered Westlaw AI-Assisted Research and Ask Practical Law AI as tested before the later agentic CoCounsel Legal relaunch. The current product architecture, branding, and workflow may differ. Treating the Stanford numbers as a measured failure rate for current CoCounsel Deep Research would be the same kind of category error lawyers complain about when an opponent cites a case for a proposition it did not decide.
At the same time, the Stanford study cannot be ignored just because the product names have shifted. It is independent evidence that leading legal research AI systems, including Thomson Reuters research products available before the later relaunch, produced materially unreliable answers under benchmark conditions [3][4]. That is a reason to demand current, task-matched evidence before relaxing verification duties.
For this evaluation, the available public record does not include an independent, preregistered benchmark of CoCounsel’s current agentic Deep Research mode. That is not proof that the mode performs poorly. It is an absence of independent evidence for the specific use that creates the highest professional exposure: generating legal research that may influence a filed brief.
The sanction record is not abstract

The reason verification cannot be treated as a polite caveat is U.S. v. Farris. In that Sixth Circuit matter, the record described a filing associated with a “CoCounsel Skill Results” filename and AI-generated material that included three fabricated quotations and misrepresented the holdings of U.S. v. Washington, 715 F.3d 975, and U.S. v. Anthony, 280 F.3d 694 [5][6].
The professional consequences were concrete. The attorney’s Criminal Justice Act compensation was denied, the matter was referred for discipline, and the attorney was removed from the case [5][6]. Those are not reputational hypotheticals. They are the kinds of outcomes that change firm policy because someone has to explain to a client, a court, and a malpractice carrier why a brief contained authorities that did not say what the lawyer claimed.
Farris is especially useful as a risk signal because it sits exactly where the benchmark evidence is weakest. It was not about whether AI can summarize a supplied document. It was about whether research output could be relied on in court. The failure mode was not subtle: fabricated quotations and misstated holdings are the parts of a brief a lawyer must verify before filing.
Fletcher v. Experian adds another CoCounsel/vLex signal. On February 18, 2026, the Fifth Circuit imposed a $2,500 sanction on attorney Heather Hersh in connection with vLex/CoCounsel use [7]. The ABA Journal also treated Fletcher as part of the broader increase in sanctions involving AI hallucinations [8]. The lesson is not that every CoCounsel output is defective. It is that courts are no longer treating unverified AI legal research as a novel mistake.
A Central District of California special-master matter from May 2025 should be handled more carefully. There, roughly 9 of 27 citations were incorrect, including two non-existent cases; the brief was struck and $31,100 in fees were awarded. But the reported workflow involved CoCounsel alongside Westlaw Precision and Google Gemini, so it should not be attributed to CoCounsel alone [9]. Its value is as a warning about mixed-tool workflows: once lawyers move output among systems, the final filing still belongs to the lawyer, not to the toolchain.
For a fuller incident register, see the internal sanctions catalog on Westlaw AI hallucination sanctions. The point here is narrower: documented cases already show the exact filing risk that a law firm must control before treating AI research as usable.
Security and scale answer different questions
Thomson Reuters has real procurement advantages. CoCounsel Legal sits inside a familiar legal-information ecosystem, connects to Westlaw and Practical Law, and advertises integrations with document and knowledge systems that large firms already use, including iManage, NetDocuments, SharePoint, HighQ, and Microsoft 365 [10]. Thomson Reuters also states that client data is not used to train underlying models and identifies SOC 2 and ISO 42001 as part of its security posture [10].
The scale is also real. Thomson Reuters said in February 2026 that CoCounsel had reached one million professionals across 107 countries [11]. The product line traces back to the March 1, 2023 launch of Casetext’s CoCounsel, Thomson Reuters’ $650 million acquisition of Casetext that closed in August 2023, and the later CoCounsel Legal relaunch [11][12].
Those facts matter in a vendor review. They help answer whether the supplier is durable, whether the tool can fit existing systems, and whether confidentiality and information-governance questions have at least been addressed in a serious way. They do not answer whether a generated citation is real, whether a quotation appears on the cited page, or whether a case still stands for the proposition asserted.
The same separation applies to pricing. Third-party reviews describe a wide 2026 pricing band for CoCounsel as a Westlaw add-on, including estimates of roughly $104 to $639 per user per month, while other reports discuss different package levels and enterprise pricing. None of those figures should be treated as an official Thomson Reuters price sheet [13][14]. Pricing opacity is a procurement issue. It is not a reliability metric.
A defensible law-firm use policy

A sensible policy does not ban CoCounsel, and it does not bless it globally. It assigns different controls to different tasks.
- For closed-document work, require the user to preserve the source set, keep the output tied to the uploaded documents, and review samples against the original record before relying on the answer.
- For litigation research, require human verification of every cited authority, quoted passage, holding, procedural posture, parenthetical, and negative-treatment status before the material is filed or sent externally.
- For mixed-tool workflows, require the final drafter to verify the final text rather than assuming that an upstream Westlaw, CoCounsel, Gemini, or document-management step preserved accuracy.
- For procurement, separate security review from accuracy review. SOC 2, ISO 42001, data-use commitments, and integrations belong in vendor diligence, but they do not replace legal-source verification.
- For training, use real sanction examples rather than abstract warnings. Lawyers remember compensation denial, referral, removal, struck briefs, and fee awards more readily than generic “AI may hallucinate” language.
A firm that wants a more formal gate can adapt the verification approach discussed in AI verification workflows for lawyers and related filing-control guidance in legal copilot setup. The important point is not the form name. It is that the lawyer who signs the filing must be able to show where each legal proposition was checked.
CoCounsel’s closed-document benchmark record is strong enough to justify serious use in controlled document-analysis workflows. Its legal-research output should remain in a different category: capable, potentially useful, and unverified until a human checks the cases, quotations, holdings, and procedural posture against primary sources before filing.
References
- CoCounsel — Application Report (VLAIR), Vals AI
- Vals Legal AI Report (VLAIR), Vals AI, February 2025
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Stanford RegLab
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies, 2025
- Hallucinations by West’s CoCounsel, EDRM, April 2026
- Sixth Circuit Sanctions Attorney for Unverified AI-Generated Briefs, Speaker Law
- Fletcher v. Experian Information Solutions, Inc., Justia, February 18, 2026
- Sanctions ramping up in cases involving AI hallucinations, ABA Journal
- AI Hallucinations Strike Again: Two More Cases Where Lawyers Face Judicial Wrath for Fake Citations, LawNext, May 2025
- CoCounsel Legal, Thomson Reuters
- One million professionals turn to CoCounsel as Thomson Reuters scales AI for regulated industries, Thomson Reuters, February 2026
- Three Years After Launching As First AI Legal Assistant, CoCounsel Reaches 1 Million Users And Thomson Reuters Teases What’s Ahead, LawNext, February 2026
- CoCounsel Review & Alternatives, HAQQ, 2026
- CoCounsel Review: Artificial Intelligence for Lawyers, Lawyerist
Chronological incident history
No sanction cases have named this tool in the tracked record set to date. This does not imply the tool is safe — see Risk Digest for ongoing monitoring.
← Compare peer toolsReport a correction or tip
Spotted an outdated figure, a misstated fact, or a ruling this tool profile should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.
Report a correction or tip for this record →