Seagate's Record Storage Shipments Expose E-Discovery Gaps
Seagate's Q4 FY2026 revenue surged 48% as it shipped 218 exabytes, mostly to AI data centers. The article explains why this storage explosion means current e-discovery preservation frameworks for AI-generated data are dangerously under-scoped, citing court rulings that already treat AI logs as discoverable.
- Jurisdiction
- US Federal
- Court
- Southern District of New York
- AI tool named
- ChatGPT
- Ruling date
- Jun 1, 2026
- Source document
- View primary court order ↗
- Last verified
- Jul 29, 2026
Lex Machina Review is an independent risk-tracking and reference resource. Nothing on this site is legal advice, and using it does not create an attorney-client relationship. Every record is reviewed against primary sources but may not reflect the most current status of a matter — always verify directly against the cited court order, rule text, or a licensed attorney before relying on it.
Companion explanation — secondary to the source document above
Seagate’s July 28 earnings call is easy to read as a market event: $3.6 billion in Q4 FY2026 revenue, up 48% year over year; 218 exabytes shipped in one quarter; 89% of those shipments going to data-center customers; and capacity effectively allocated through calendar 2028, with customers already asking about 2029 visibility.[1] For discovery counsel, that is the less interesting version of the story. The sharper signal is physical: AI systems are generating records at a scale large enough to consume hard-drive capacity years before it exists.
That is where Seagate’s AI storage demand outlook crosses into litigation risk. The point is not that every additional exabyte becomes a preservation obligation. It is that the enterprise systems now being built for AI are making prompts, outputs, context stores, tool traces, and conversation histories less like passing remarks and more like stored business records.

The Storage Number Is the First Warning
A quarterly shipment figure like 218 exabytes does not prove a preservation crisis by itself. Seagate sells into many workloads, and an earnings-call transcript is not a court record or an audited legal-risk study. But the mix matters. On the call, Seagate tied the surge to nearline demand from data centers, especially AI infrastructure, and said data-center customers represented 89% of capacity shipments in the quarter.[1]
The more important detail came from CEO Dave Mosley’s explanation of why AI storage demand is changing. He described agentic AI workflows and KV cache architecture as a reason data is “not only growing, it is compounding,” because persistent context is stored on hard-drive tiers rather than disappearing after a single inference session.[1] That phrase deserves more attention from lawyers than from traders.
Traditional discovery planning is still often organized around familiar containers: email, shared drives, collaboration platforms, document repositories, databases, mobile messages. Those categories remain essential. They also do not describe the full record trail created when AI agents retain context, call tools, write intermediate outputs, preserve session memory, and feed later interactions with earlier ones.
Why Agentic AI Logs Are Not Just Longer Chat Histories
A one-off chatbot exchange can look deceptively simple: a user enters a prompt, the system returns an answer, and the visible exchange ends. Agentic AI changes the discovery analysis because the useful business record may sit around the visible exchange rather than only inside it.

An agent may receive a prompt, retrieve internal documents, call a sales database, summarize a customer record, draft a recommendation, update a task list, and retain parts of that interaction so the next session begins with more context. In that setting, the discoverable candidate is not only the final answer. It may include the prompt, the output, retrieved materials, tool calls, intermediate reasoning artifacts if stored, system instructions, user edits, deleted messages, session metadata, and whatever persistent context the platform keeps to make the next interaction more useful.
KV cache is a technical term, not a litigation doctrine. Still, Mosley’s description matters because it points to the persistence layer. If AI infrastructure is being designed so prior context can be stored, reused, and accumulated, then counsel cannot safely assume that an AI interaction is ephemeral merely because it resembles chat in the user interface.[1]
| AI artifact | Discovery significance |
|---|---|
| User prompts | May show what the user asked the system to do, including knowledge, intent, scope, or instructions. |
| Model outputs | May show generated recommendations, summaries, drafts, or representations relied on by a person or business process. |
| Conversation logs | May preserve the sequence of questions, corrections, deletions, and refinements. |
| Tool calls and retrieval traces | May identify source systems consulted by the AI agent and the documents or data used. |
| Persistent context stores | May carry forward facts, preferences, or summaries from one session into later conduct. |
The table is not a preservation rule. It is a map of places where records may exist. Whether any item must be preserved depends on relevance, proportionality, control, foreseeability, and the particular system configuration. The practical problem is that many litigation-hold templates do not ask the threshold questions at all.
Courts Are Already Treating AI Records as ESI
The legal signal is still developing, and some of the available public discussion comes through vendor analysis rather than primary docket materials. That caveat matters. But the reported rulings are already enough to make “AI data is different” a weak preservation assumption.
In June 2026, Smarsh described In re OpenAI, Inc. Copyright Infringement Litigation in the Southern District of New York as compelling production of millions of ChatGPT conversation logs, treating AI-generated conversation data as discoverable electronically stored information under existing discovery rules.[2] Before relying on that description in a brief, counsel should verify the underlying order against the docket. For planning purposes, however, the reported direction is plain: courts do not need a new category called “AI evidence” before they can order production of AI records.
Smarsh also reported a March 2026 Delaware Court of Chancery ruling in a $250 million earnout dispute that admitted a CEO’s ChatGPT logs as evidence of intent and treated deleted messages as a spoliation concern.[2] That is a narrower point than saying all AI chats will be decisive. It is also more useful. When the contents of an AI exchange bear on intent, knowledge, drafting history, valuation assumptions, or post-dispute conduct, the log can move from technical exhaust to merits evidence.
The sanctions examples point to a related but distinct failure mode. Smarsh identified Whiting v. City of Athens in the Sixth Circuit as involving a $30,000 sanction for AI-generated fabricated citations, and State v. Gorso in Massachusetts as involving a $10,000 sanction for AI-generated fabricated citations, both within the past year.[2] Those cases are not storage-volume cases. They show that courts are already attaching monetary consequences to careless handling of AI outputs in litigation work.
The Preservation Gap Is a Mismatch, Not a Proven Industry Statistic
There is no need to overstate the record. The available sources do not establish that most companies have defective AI-retention policies, or that law firms are uniformly ignoring AI artifacts. The preservation gap is a disciplined inference from two facts now moving toward each other: AI infrastructure is being built for persistent, compounding data, and courts are already willing to treat AI conversations and outputs as discoverable ESI.
The weak point in many holds is naming. A hold that instructs custodians to preserve “documents, emails, text messages, spreadsheets, presentations, databases, and collaboration-platform messages” may sound comprehensive while still failing to identify the AI systems where relevant conduct occurred. If a sales team used an AI assistant to generate account strategy, if an executive used ChatGPT to test deal language, or if an engineering group used an agent to summarize defect reports, the hold has to reach the places those interactions were stored.
The hard question is not whether a company should preserve every AI artifact forever. That would be an impractical and legally imprecise answer. The harder question is whether the organization can identify which AI tools are in use, who controls their logs, what default deletion settings apply, whether deleted conversations remain recoverable, whether prompts and outputs are exported into other systems, and whether agentic context stores can be suspended, preserved, searched, or defensibly excluded.
The Capacity Curve Looks Structural
Seagate’s quarter would be less significant if it looked like a one-quarter inventory event. The company’s capacity comments argue against treating it that way, though earnings-call statements still need cross-checking against formal filings for contractual precision. On the call, Seagate said capacity was effectively allocated through 2028 and that customers were seeking visibility into 2029.[1]
The product roadmap points in the same direction. Seagate has said its Mozaic 4+ platform, with 44TB drives and more than 4TB per disk, is shipping in volume to two hyperscale cloud providers, while Mozaic 5+ qualification shipments are expected in late calendar 2027.[3] That is not legal evidence of preservation failure. It is infrastructure evidence that the storage layer beneath AI systems is being expanded deliberately, not accidentally.
A Forbes piece by Tom Coughlin, using a Coughlin Associates estimate, projected 363 exabytes of additional HDD capacity demand from AI in 2026, or about 18% of all capacity shipments, rising to 43% by 2028 and 58% by 2030.[4] That is one analyst firm’s model, not an audited industry count. Used carefully, it reinforces the same limited point: AI storage demand is expected to take a larger share of HDD capacity over the next several years.
What the Next Litigation Hold Has to Name
A useful AI-aware hold does not need to become a technical manual. It does need to stop treating AI as a vague productivity tool. The preservation notice, custodian questionnaire, and data-map interview should be concrete enough to identify systems and artifacts before auto-deletion, vendor retention settings, or ordinary model-workflow churn make the answer unrecoverable.
- Which generative AI and agentic AI tools did custodians use for the relevant business activity?
- Are prompts, outputs, conversation histories, uploaded files, retrieved materials, and tool-call logs stored?
- Can users delete AI conversations, and if so, are deleted messages retained elsewhere?
- Do persistent context, memory, or KV-cache-related stores carry information from one session into another?
- Which vendor, cloud tenant, workspace, or internal system controls export, search, retention, and suspension?
Those questions are modest compared with the infrastructure buildout they are trying to track. Seagate’s record quarter does not mean every AI log is relevant, discoverable, or proportionate in every dispute. It does mean that the storage backbone for AI-generated records is no longer theoretical. If capacity is already being allocated through 2028, the next litigation hold should be able to say what it means by prompts, outputs, conversation logs, deleted AI messages, and agentic context stores.
References
- Seagate Technology Fiscal Q4 FY2026 Earnings Call Transcript, MarketBeat, July 28, 2026, link
- Smarsh June 2026 analysis of AI records, discovery rulings, and sanctions, Smarsh, June 2026, link
- Mozaic 4+ announcement, Seagate Investor Relations, link
- Forbes article by Tom Coughlin on AI HDD capacity demand, Forbes, May 30, 2026, link
Related records
Tool profile
How to Read Legal AI Benchmarks Ahead of the OpenAI IPOGoverning regulation
Browse the obligations tracker →Preventive workflow
Browse verification workflows →
Report a correction or tip
Spotted an outdated figure, a misstated fact, or a ruling this case record should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.
Report a correction or tip for this record →