Vetting AI Training Data After Thomson Reuters v. ROSS
Thomson Reuters v. ROSS shows why training-data provenance, not output behavior, is the decisive procurement risk for legal-AI tools. Use the included vendor-diligence checklist — source attestation, chain of title, clean-room evidence, audit rights, and indemnities — when vetting or renewing any tool trained on editorial content.
- Tool
- ROSS Intelligence
- Benchmark source
- District of Delaware summary-judgment ruling
- Hallucination rate
- Not measured / undisclosed
- Test methodology
- Fair-use factor analysis of summary-judgment record
- Test date
- Aug 26, 2026
The reassuring sentence that should slow the purchase
A legal-AI vendor can say three things that sound calming in a procurement meeting: the tool is non-generative, it does not show copied source text to users, and any disputed material was only a tiny part of a much larger corpus. The Thomson Reuters v. ROSS Intelligence copyright lawsuit over AI training data is uncomfortable because those statements did not answer the risk that mattered. In February 2025, the District of Delaware held that ROSS infringed Westlaw headnotes used in training materials even though the legal-research tool did not display those headnotes as outputs.[1]
That makes the case more useful as a diligence file than as another abstract fair-use debate. For the full chronology, the procedural posture, and the running appeal tracker, use the existing case-status explainer. The procurement lesson is narrower and more urgent: output behavior does not prove training-data provenance.

Where the provenance problem entered the record
The record did not start with a chatbot generating a memorized passage. It started with licensing and sourcing. Thomson Reuters refused to license Westlaw content to ROSS. ROSS then commissioned roughly 25,000 “Bulk Memos” from LegalEase. Those memos were derived from Westlaw headnotes, despite instructions not to copy headnotes directly. On summary judgment, the court held that 2,243 of 2,830 asserted headnotes were infringed “such that no reasonable jury could find otherwise.”[1]
For a buyer, the warning sign is not that a contractor was involved. Contractors are routine. The warning sign is the gap between the instruction and the proof. “Do not copy” is a control only if the vendor can show what was actually used, who reviewed it, how derivative material was screened, and what records survived.
| Vendor assurance | What the assurance leaves unanswered | Why ROSS matters |
|---|---|---|
| “The tool is non-generative.” | What materials were used to train or tune the system? | ROSS was a legal-research system returning judicial opinions, not a generative writing product, yet the training materials drove liability risk.[1] |
| “Users never see copied source text.” | Was protected editorial material copied upstream to build the model or training examples? | The district court treated training-copying as material even without copied headnotes appearing in outputs.[1] |
| “Only a small fraction of the source library was implicated.” | Was the copied material qualitatively important, curated, or market-substituting? | At appeal, ROSS emphasized a 0.08% figure; Thomson Reuters countered that the material captured the heart of its editorial work.[2] |
| “Our contractors were told not to copy.” | Can the vendor produce clean-room records, review logs, sampling results, or independent-analysis evidence? | The ROSS record turned on what the Bulk Memos actually contained, not on the comfort of the instruction.[1] |
Why output independence did not carry the fair-use analysis
The district court’s factor-one analysis treated ROSS’s use as commercial and not transformative. Applying the Supreme Court’s Warhol framing, the court focused on whether ROSS used the headnotes for a purpose sufficiently distinct from Westlaw’s legal-research function. It concluded that ROSS used the headnotes to build a competing legal-research tool, not for a meaningfully different purpose.[1]
Factor two helped ROSS but did not save it. The court described the works as thin, because headnotes sit close to judicial opinions and legal propositions. That mattered, but it did not overcome the court’s findings on commercial purpose and market harm.[1]
Factor four is the part that belongs in procurement review. The court credited harm to an actual or potential market, including an AI-training licensing market, even though Thomson Reuters had not already established such a market. In other words, a vendor cannot assume that the absence of an existing license category means there is no licensing-market risk.[1]
That is the practical edge of the ruling. If a product competes with the source owner’s product, and if the disputed training material comes from curated editorial work rather than public-domain source law, the buyer should not treat “no copied output” as the end of the copyright inquiry.
The small-percentage argument is a weak procurement control
At the Third Circuit argument on June 11, 2026, ROSS pressed the point that the disputed headnotes represented about 0.08% of Westlaw’s 28-million-headnote library. Thomson Reuters answered that ROSS had taken the “heart” of Westlaw’s editorial content for the task ROSS needed, and LawNext reported that ROSS’s own expert had conceded ROSS could have built training data from the public opinions themselves.[2]
A tiny percentage can still be a serious procurement risk when the percentage is calculated against the wrong denominator. The denominator a risk committee cares about is not only the source owner’s full archive. It is also the subset used to create the competing capability, the materiality of the copied selection, and whether the vendor could have obtained equivalent training examples from rights-cleared sources.
Public-law output does not clean proprietary input
ROSS returned judicial opinions, which are public law. Westlaw headnotes are different. They are editorial enhancements layered onto public opinions. That distinction is the reason chain of title matters so much in legal-AI diligence. A vendor may have the right to use court opinions and still lack rights to the annotations, summaries, classifications, headnotes, or taxonomy attached to them.
The appeal may change the legal analysis. The Third Circuit heard argument in June 2026, and Baker Botts reported in July 2026 that a decision was expected in late 2026.[3] As of Aug. 26, 2026, there is no appellate decision and no damages trial. The district ruling is a current risk signal, not the final word.
The diligence file a buyer should require
This checklist is the copyright-and-provenance counterpart to a broader legal-AI vendor due-diligence checklist. It is aimed at the part of the file that too often arrives late: what the vendor trained on, who had rights to it, and what happens if that answer proves wrong. Practitioner guidance after ROSS has converged around the same core controls: source attestation, chain of title, clean-room evidence, audit rights, warranties, indemnities, litigation disclosure, and, where relevant, EU AI Act transparency obligations.[4][5][6]

| Control | What to request before approval | ROSS-linked reason |
|---|---|---|
| Training-data source attestation | A signed statement identifying training, fine-tuning, evaluation, and retrieval corpora by source category, acquisition path, date range, and material exclusions. The attestation should distinguish public judicial opinions from editorial enhancements. | The disputed risk in ROSS sat upstream in training materials, not in the final user-facing output.[1] |
| Chain of title for editorial content | License schedules or ownership records showing rights to headnotes, annotations, summaries, taxonomies, classification systems, and other editorial layers. General permission to access legal materials is not enough. | Westlaw headnotes were treated differently from the public opinions they summarized.[1] |
| Contractor and subprocessor provenance | Names and roles of any contractors who created training examples, memos, labels, summaries, or evaluation sets; copies of relevant statements of work; and records showing how deliverables were reviewed. | ROSS’s Bulk Memos were contractor-created materials, and the court examined what they contained rather than accepting instructions at face value.[1] |
| Clean-room or independent-analysis evidence | Documentation of separation from proprietary sources, reviewer instructions, similarity screening, sampling results, escalation logs, and remediation steps for any contaminated examples. | A policy against copying does not establish that no protected expression entered the dataset. |
| Audit rights | A contractual right, exercisable under confidentiality limits, to review training-data records, provenance manifests, contractor files, and material changes to training sources. If direct inspection is impossible, require a third-party audit report with scope defined in the contract. | A buyer needs a way to test provenance representations before a claim arrives, not only after litigation begins. |
| Provenance warranties | Specific warranties that training and evaluation data were lawfully obtained and used, that the vendor has rights sufficient for the contracted use, and that no undisclosed third-party restrictions impair deployment. | A generic “we comply with law” clause is too soft for a dispute centered on an upstream license refusal and later workaround. |
| Copyright indemnity and operational remedies | Defense and indemnity for copyright claims tied to training data, plus remedies for injunction risk: replacement model, retraining obligation, service continuity plan, refund rights, and termination assistance. | The practical loss to the buyer is not only damages exposure. It is interruption, forced migration, and a scrambled knowledge workflow. |
| Pending-litigation and demand-letter disclosure | Disclosure of pending copyright cases, threatened claims, indemnity tenders, license disputes, takedown demands, or regulatory inquiries involving training data or source materials. | A renewal decision under time pressure should not depend on whether the buyer happened to find the vendor’s litigation history. |
| Change control for retraining and new corpora | Notice and approval rights for adding new proprietary legal corpora, contractor-generated examples, or third-party datasets after signing. Tie each material model update to a dated provenance record. | A clean diligence file at procurement can become stale after retraining. |
| GPAI and transparency obligations where relevant | For deployments touching EU AI Act obligations, request the vendor’s training-data summary approach, documentation process, and allocation of transparency responsibilities. | Transparency rules do not replace copyright rights analysis, but they can expose whether the vendor has a disciplined source inventory.[6] |
The request should be made before the security questionnaire is treated as complete. Security review can tell a buyer where data goes during use. It usually does not tell the buyer where the model came from.
Use tool class as a risk signal, not a shortcut
Not every legal-AI product presents the same provenance profile. A comparison by practice area and workflow belongs in a broader legal-research tool-class review, but for copyright diligence the first sort is simpler: what source material gives the tool its legal knowledge?
| Tool class | Typical provenance question | Procurement posture |
|---|---|---|
| Licensed-corpus retrieval or research tools | Does the vendor’s license cover the specific use: search, retrieval, ranking, training, fine-tuning, evaluation, and customer deployment? | Lower concern if license scope is explicit; do not assume a content-access license covers model training. |
| Tools trained on proprietary editorial legal content | Did the vendor use headnotes, annotations, citators, summaries, classification systems, or expert-created labels from a publisher or platform? | Highest ROSS-like scrutiny. Require chain of title, clean-room records, and indemnity before approval. |
| Tools trained primarily on public legal materials | Were public opinions, statutes, regulations, and filings separated from proprietary editorial overlays and commercial datasets? | Public source status helps only if the vendor can show that proprietary layers were not imported with the public material. |
| Broader generative models using large public or web corpora | What copyrighted material was included, what licenses or exceptions are claimed, and what litigation or policy changes could affect deployment? | Different exposure profile. Output memorization controls matter, but they do not replace training-data provenance review. |
This is also where the broader AI-training fair-use split matters. District courts have not moved in one direction. Readers tracking Bartz, Kadrey, and ROSS together should use the site’s fair-use divide overview and district-ruling comparison. For procurement, the unsettled law cuts against relying on vendor confidence statements. It does not make diligence optional.
How to make the checklist enforceable
A provenance checklist is useful only if it changes the contract file. The vendor’s answers should not sit in a slide deck outside the agreement. Attach the source attestation to the order form or master agreement. Define material deviations as a breach or at least as a notice-and-approval event. If the vendor later retrains on a new legal corpus, the buyer should not have to rediscover the provenance issue during the next renewal.
- Separate output controls from input rights. A non-reproduction warranty addresses one risk; it does not warrant lawful training data.
- Require named exclusions when the vendor says it did not use proprietary legal editorial content. “We use public legal materials” should be followed by what was excluded.
- Escalate if the vendor used contractor-created legal summaries, labels, memos, or examples. Ask for the contractor workflow, not just the vendor’s top-level policy.
- Make indemnity operational. Defense costs are not enough if the remedy leaves lawyers without the research tool mid-matter.
- Calendar appeal-driven re-review. A clause requiring legal-change notice is better than hoping someone remembers to revisit the file.
The buyer also needs an internal owner. Someone in legal ops, KM, procurement, or the GC’s office should be responsible for keeping the attestation, license schedule, audit report, indemnity language, and litigation disclosure together. Fragmented diligence is how a rights issue becomes invisible.
What remains provisional
As of Aug. 26, 2026, the district ruling in Thomson Reuters v. ROSS makes training-data provenance a current procurement risk for legal-AI tools. It does not settle AI training law nationally. The appeal remains pending in the Third Circuit after the June 11, 2026 argument, the district ruling is not binding outside its court, and damages have not been tried.[3]
The right file today is therefore conditional: approve, renew, or reject based on the present record, but mark the provenance checklist for re-verification when the appellate decision lands. For monitoring sources tied to that update cycle, keep a live watchlist of AI copyright and fair-use research tools.
References
- Memorandum Opinion, Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 1:20-cv-00613-SB — U.S. District Court for the District of Delaware — Feb. 11, 2025
- At 3rd Circuit, Judges Press ROSS and Thomson Reuters on Fair Use, AI Training and Market Harm — LawNext — June 2026
- Third Circuit Hears Oral Argument — Baker Botts — July 2026
- Court AI Fair Use Thomson Reuters Enterprise GmbH ROSS Intelligence — Reed Smith — Mar. 3, 2025
- Reuters v ROSS Court Ruling AI Copyright Fair Use — Davis Wright Tremaine — Feb. 2025
- AI in litigation series: An update on AI copyright cases in 2026 — Norton Rose Fulbright — 2026
Chronological incident history
No sanction cases have named this tool in the tracked record set to date. This does not imply the tool is safe — see Risk Digest for ongoing monitoring.
← Compare peer toolsReport a correction or tip
Spotted an outdated figure, a misstated fact, or a ruling this tool profile should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.
Report a correction or tip for this record →