Skip to content

Risk Digest

Three Due Diligence Gaps Emerging from PE AI Investment Trends

The $63B PE-led AI investment wave in 2025 introduced novel legal risks — data provenance gaps, shadow AI exposure, and regulatory compliance costs — that standard software diligence frameworks fail to address. This analysis maps those risks and the evolving representations, warranties, and RWI exclusions that deal lawyers must navigate.

By Editorial TeamUpdated Jul 25, 2026Verified Jul 25, 2026
REPORTED — UNVERIFIED
Jurisdiction
United States
Court
U.S. District Court
AI tool named
Generative AI
Ruling date
Dec 31, 2025
Source document
View primary court order ↗
Last verified
Jul 25, 2026

Lex Machina Review is an independent risk-tracking and reference resource. Nothing on this site is legal advice, and using it does not create an attorney-client relationship. Every record is reviewed against primary sources but may not reflect the most current status of a matter — always verify directly against the cited court order, rule text, or a licensed attorney before relying on it.

Companion explanation — secondary to the source document above

Private equity did not wait for AI diligence practices to mature. In 2025, Crunchbase counted $63 billion in PE-led AI investment and $202.3 billion in total AI funding, while CB Insights reported a broader $225.8 billion AI funding total, with 79% of that funding tied to mega-rounds and 38% flowing to three frontier labs.[1][2] The exact market total depends on methodology, but the legal point does not: capital moved faster than many targets’ ability to prove what they owned, what their employees used, and what regulatory posture a buyer would actually inherit.

Deal count tells the same story with less precision. WithIntelligence reported 589 PE AI deals in 2025, up 57% year over year, though direct verification of the underlying page was limited and should be treated accordingly.[3] Even with that caveat, the direction is hard to miss. AI moved from a specialist investment category into ordinary sponsor deal flow, and ordinary software diligence began showing its seams.

Deal room documents with diligence gaps beneath a hovering AI data network

The question for lawyers is not whether AI companies attracted record attention. It is what that wave revealed that a standard code, cyber, IP, and vendor review does not catch. Three gaps now deserve their own diligence workstreams: undocumented data provenance and model ownership, hidden employee use of generative AI tools, and a regulatory overlay that can change cost, remedies, insurance, and exit value.

The first gap is not code ownership. It is the chain of rights behind the model.

Traditional software diligence has a familiar rhythm: confirm who wrote the code, identify open-source exposure, review inbound and outbound licenses, test assignment agreements, and map material vendors. That still matters. It is just no longer enough when the asset being priced is a model trained, tuned, evaluated, or improved on data that may have been licensed, scraped, purchased, contributed by customers, generated by users, or created inside a corporate workflow.

Mayer Brown’s May 2026 analysis puts the point in deal-room terms: for AI targets, the absence of documentation is itself a material finding, not a neutral answer waiting for cleanup after signing.[4] That is the sentence that should change the diligence request list. If the target cannot show where training data came from, what rights attached to it, whether those rights permit model training or commercial deployment, and whether any customer or employee data was used outside the permitted scope, the buyer is not merely missing a schedule exhibit. It may be missing the proof needed to own or exploit the asset it thinks it is buying.

Broken data rights chain connecting data sources to a central AI model

Morgan Lewis described data provenance, model IP, compute access, and explainability as central diligence and negotiation points in 2025 AI deals.[5] Those are not four decorative subtopics. They are different ways of asking whether the buyer can verify the economic asset. Data provenance asks whether the model was built on usable inputs. Model IP asks who owns the resulting weights, fine-tunes, embeddings, prompts, outputs, evaluation sets, and related improvements. Compute access asks whether the target can continue operating the product on commercially viable terms. Explainability asks whether the target can support customer, regulator, and contractual demands when an AI system produces a consequential output.

The hard cases are not always dramatic. A target may have a neat product demo and a competent engineering team, but only informal records showing that a model was trained on mixed data: some licensed content, some customer-supplied material, some public web data, and some internal support tickets. Each category raises a different question. Did the license permit training? Did the customer contract allow secondary use? Was the public data subject to terms that restricted automated collection or commercial exploitation? Were support tickets anonymized, and was anonymization enough under the relevant privacy and confidentiality obligations?

That is where legal ownership and deal value collide. A model may be commercially impressive and still be legally awkward. A buyer can price known remediation, negotiate a special indemnity, exclude a data set, require retraining, or demand a closing deliverable. What it cannot do sensibly is treat an undocumented rights chain as if it were a missing employee invention agreement from a small subsidiary. The defect may sit inside the model’s performance, not beside it.

Diligence questionWhy the answer changes the deal
What data trained, tuned, evaluated, or improved the model?Identifies whether the target can prove lawful and contractually permitted use of the inputs.
Who owns the model, weights, fine-tunes, embeddings, prompts, outputs, and evaluation assets?Tests whether the buyer is acquiring the core AI asset or only an operating wrapper around disputed components.
Which customer, employee, or third-party data entered the system?Surfaces confidentiality, privacy, contractual-use, and consent issues before they become post-closing claims.
What compute, API, or infrastructure dependencies are required to run the product?Separates ownership of the model from the practical ability to operate it at the forecast margin.
Can the target explain model behavior where customers or regulators require it?Affects regulated deployments, customer contracting, and the cost of future compliance.

A good AI diligence request therefore should not ask only for “all IP licenses” and “all open-source software.” It should ask for the training-data inventory, data acquisition records, terms governing each material data source, consent and usage limitations, customer-contract provisions governing secondary use, records of model development and fine-tuning, evaluation and red-team materials, model cards or similar documentation where available, and any internal approvals for using personal, confidential, or regulated data. If the target has none of this, that answer belongs in the risk memo, not in a footnote to be chased later.

Shadow AI is the risk the vendor list will not show

The second gap is more irritating because it can hide in plain sight. Shadow AI is not the target’s flagship model or a disclosed enterprise AI vendor. It is the engineer pasting logs into a public chatbot, the sales team using a generative tool to draft customer responses, the HR manager summarizing applicant materials, or the finance team uploading portfolio data into an unapproved assistant. Mayer Brown identifies shadow AI as a distinct PE deal risk because informal use of public generative AI tools can create data leakage, IP, confidentiality, and contractual exposure that ordinary diligence may miss.[4]

Hidden unapproved AI tools leaking data beneath visible corporate IT systems

A standard software review is poorly designed for this. The IT inventory may list approved systems. The vendor schedule may show material contracts. The security questionnaire may ask about access controls, encryption, incident response, and third-party processors. None of that necessarily captures employees using free or individually expensed AI tools outside procurement, outside security review, and outside the company’s data-processing map.

The legal consequences are not theoretical simply because the tool was informal. Confidential information can leave the corporate perimeter. Customer data may be processed by an undisclosed provider. Contract restrictions on data use may be breached without anyone in legal seeing the workflow. Employee-created prompts or outputs may become part of product development without a clean record of source material. If regulated or sensitive information enters a public tool, the buyer may inherit both a disclosure problem and an operational habit.

This is why “no written AI policy” is not a neutral answer. It means the buyer has to assume the control environment may be undocumented. The target might still have good practices: enterprise licenses, disabled training by the provider, DLP controls, logging, employee training, and legal review for sensitive use cases. But those facts need evidence. A clean slide saying the company is “responsible AI by design” should not survive contact with an empty policy folder.

The better diligence questions move from inventory to behavior. Which AI tools are blocked, allowed, or monitored? Are employees permitted to input source code, customer data, personal information, trade secrets, deal materials, or privileged communications? Does the company log prompts or only approved enterprise-tool usage? Have any teams adopted AI tools through personal accounts or browser extensions? Has the company reviewed outputs before using them in product code, marketing claims, customer deliverables, hiring, credit, insurance, or other consequential decisions?

A buyer should also expect imperfect answers. Shadow AI is partly a discovery problem. Interviews with engineering, sales, customer support, HR, finance, and legal may matter as much as the vendor list. Expense reports, browser-extension policies, SSO logs, endpoint controls, DLP alerts, and procurement exceptions can tell a different story from the formal questionnaire. The point is not to turn every deal into a forensic investigation. It is to know when the target’s AI use is governed, observable, and contractually consistent, and when it is merely convenient.

Regulation now affects price, not just post-closing compliance

The regulatory overlay should be handled with discipline. Not every AI target is a high-risk system, and not every rule is fully in force as of Q3 2026. But sponsors cannot treat AI regulation as a distant public-policy discussion when it affects compliance spend, customer contracting, product design, and exit value.

Skadden warned in January 2025 that rising AI investment required financial sponsors to address unique risks, including EU AI Act compliance costs and Colorado AI Act bias-audit obligations.[6] For the EU AI Act, the highest-friction question is whether the target’s system falls into a category that triggers high-risk obligations. Those obligations take full effect on August 2, 2026, so they should be treated as a near-term compliance and cost issue rather than a settled 2025 operating fact.[6] Readers tracking the timeline can use the internal guide to the EU AI Act high-risk deadline and the broader 2026 AI compliance calendar.

Colorado requires similar care. The available support is narrower: the Colorado AI Act remains relevant to consequential-decision systems, including employment and financial services, while legal services were subsequently removed from the state’s AI regulation. That distinction matters for law-firm readers, but it does not make the statute irrelevant to PE portfolios with hiring, lending, insurance, housing, education, or other covered use cases. For state-level context, see the internal updates on state AI laws affecting law firms in 2026 and Colorado’s legal-services exclusion.[6]

Enforcement risk also enters through marketing. Mayer Brown links AI deal risk to SEC and FTC scrutiny of AI-related claims, including the risk that companies overstate capabilities, controls, or compliance posture.[4] That changes the diligence review of customer decks, website claims, investor materials, security questionnaires, SOC-related statements, and regulatory mappings. A buyer needs to know not only what the product does, but what the target has said it does. The gap between the two may become a purchase-price issue, an indemnity issue, or an exit diligence issue for the next buyer.

The practical work is to map regulation to use cases. A target selling developer tooling raises different questions from a target supporting hiring, clinical triage, credit screening, fraud detection, insurance pricing, or legal-document review. The diligence team should identify where AI outputs affect individuals, where customers rely on those outputs for regulated decisions, where the target contractually promises compliance, and where remediation would require product redesign rather than a policy update. Readers evaluating tooling for that exercise may find the internal AI compliance software buyer’s guide useful after the legal mapping is complete.

The diligence gaps are now drafting problems

Once these risks are found, they do not sit politely in the diligence report. They move into representations, warranties, covenants, schedules, indemnities, RWI underwriting, and sometimes the deal structure itself.

A generic IP representation that the target owns or has the right to use its software may not adequately cover training data, model weights, fine-tuned models, embeddings, prompts, synthetic data, evaluation sets, and outputs. A privacy representation may not capture employee prompting of public AI tools. A compliance representation may not address whether the target has mapped AI systems to EU, state, sectoral, or customer-specific obligations. A no-claims representation may say little about a future customer dispute over AI-generated output if the underlying facts are already sitting in an undisclosed support ticket or sales escalation.

  • Representations should address training-data rights, model ownership, permitted data use, AI-related customer commitments, and accuracy of AI capability and compliance claims.
  • Disclosure schedules should identify material models, data sources, third-party AI tools, customer restrictions, known unauthorized AI use, and regulatory classifications or assessments.
  • Covenants should control pre-closing changes to model training, data ingestion, customer deployments, public AI claims, and use of unapproved generative AI tools.
  • Indemnities may need to isolate known provenance defects, customer-data misuse, specific regulatory remediation, or claims tied to particular models or data sets.
  • Closing deliverables may include AI policies, data maps, customer consents, vendor amendments, model documentation, or evidence of enterprise controls for employee AI use.

Insurance is another place where optimism meets underwriting. Mayer Brown notes that RWI carriers are introducing AI-specific exclusions.[4] That should matter early in the process, not after the purchase agreement is nearly settled. If a carrier excludes AI-related IP, data provenance, regulatory, or product-performance matters, the buyer has to decide whether to live with the exposure, reallocate it to the seller, reduce price, require remediation, or walk away from a use case that cannot be diligenced.

The carve-out problem is especially sharp when models, data sets, infrastructure, or AI personnel are shared across a seller’s retained business and the target being acquired. Ordinary transition services can provide payroll, finance support, IT access, and back-office continuity. They may be far less useful when the target’s product depends on a pooled training corpus, a shared foundation-model instance, a common prompt library, a centralized ML platform, or engineers who support multiple business units. A TSA can give access for a period of time. It cannot easily cure the absence of separable rights.

That drafting reality should feed back into diligence. If the buyer cannot separate the acquired model from retained data, cannot obtain rights to continue training, cannot verify whether customer data was used consistently with contracts, or cannot replace a shared compute arrangement without impairing performance, the issue is structural. It may require a pre-closing restructuring, a license with audit and use rights, a clean-room retraining plan, a special indemnity, or a different valuation assumption.

What changes in the diligence workstream

The most useful change is not a longer software checklist. It is a separate AI diligence workstream with legal, technical, privacy, security, product, and regulatory inputs. The workstream should begin before exclusivity hardens the timetable, because the most important findings often require interviews and document reconstruction, not just contract review.

WorkstreamMinimum output
Data provenance and model IPA rights map connecting material data sources, model assets, development history, ownership claims, and usage restrictions.
Shadow AI and employee controlsA policy-and-behavior review showing approved tools, prohibited inputs, monitoring, training, exceptions, and evidence of informal use.
Regulatory and customer obligationsA use-case map identifying high-risk or consequential deployments, customer commitments, compliance gaps, and remediation cost.
RWI and drafting allocationA coverage readout showing likely exclusions, enhanced reps, special indemnities, covenants, and schedule disclosures.
Carve-out readinessA separability analysis for shared models, data sets, compute, personnel, licenses, and transition-service limits.

The associate building the request list should ask for evidence in the form it actually exists: data maps, model cards, governance committee minutes, security exceptions, procurement tickets, customer-contract matrices, DLP logs, training materials, prompt-use policies, vendor configurations, and product-risk assessments. If the answer is that the company has operated responsibly but never documented the basis for that statement, the diligence memo should say so plainly.

The KM lawyer updating the firm checklist should resist adding a decorative AI appendix that asks whether the target “uses AI.” Almost every relevant target will say yes, and the answer will not allocate risk. Better questions ask what AI systems are material, what data they use, who owns them, what claims have been made about them, what laws and customer obligations govern them, what controls exist around employee use, and what would break if disputed data or shared infrastructure had to be removed.

The PE in-house lawyer has the harder internal conversation. Investment colleagues may accept uncertainty when the commercial case is strong. But some AI uncertainty is not just a forecasting issue. If provenance cannot be documented, hidden usage cannot be scoped, regulatory exposure cannot be mapped, and RWI will not respond, the buyer is accepting a risk that may be difficult to price and harder to carve out after closing.

The 2025 PE AI investment wave does not show that AI deals are too risky to do. It shows that AI deals are too specific to diligence as ordinary software transactions with a brighter growth curve. The closing package now has to verify the asset, expose hidden use, allocate regulatory cost, and test insurability. Where it cannot, the buyer should know that before signing, not when the first subpoena, customer claim, or renewal diligence request arrives.

References

  1. 6 Charts That Show The Big AI Funding Trends Of 2025, Crunchbase, 2025.
  2. State of AI 2025 Report, CB Insights, 2025.
  3. Private Equity AI Deals: Share of Transactions Nearly Doubles, WithIntelligence.
  4. AI: The Next Frontier of PE Deal Risk, Mayer Brown, May 2026.
  5. AI Deals in 2025: Key Trends in M&A, Private Equity, and Venture Capital, Morgan Lewis, September 2025.
  6. Rising Investment in AI Requires Financial Sponsors To Address Unique Risks, Skadden, January 2025.

Report a correction or tip

Spotted an outdated figure, a misstated fact, or a ruling this case record should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.

Report a correction or tip for this record →