Skip to content

Risk Digest

Why a ChatGPT Family Reunion Is a Privacy Warning for Lawyers

The viral story of ChatGPT reuniting a man with his half-sister by reproducing a personal blog post from training data illustrates the same memorization mechanism that exposes client confidences and raises privilege-waiver risks for legal professionals using the tool.

By Editorial TeamUpdated Jul 26, 2026Verified Jul 26, 2026
REPORTED — UNVERIFIED
Jurisdiction
jurisdiction-eu
Court
General
AI tool named
ChatGPT
Ruling date
Jul 25, 2026
Source document
View primary court order ↗
Last verified
Jul 26, 2026

Lex Machina Review is an independent risk-tracking and reference resource. Nothing on this site is legal advice, and using it does not create an attorney-client relationship. Every record is reviewed against primary sources but may not reflect the most current status of a matter — always verify directly against the cited court order, rule text, or a licensed attorney before relying on it.

Companion explanation — secondary to the source document above

The most revealing detail in the story of ChatGPT helping Avtar Singh find a lost family member is not that Singh found Nicci Dhamu. It is how he found her. Singh, who had avoided DNA testing and had deactivated Facebook because of privacy concerns, asked ChatGPT for help locating a half-sister he had never met. The model returned text that pointed him to Dhamu’s personal blog article, including material about her family background. The result was a reunion that readers could fairly find moving. It was also a privacy event: the connection depended on ChatGPT reproducing material that had been absorbed into its training data rather than disclosed by the person being searched for at that moment.[1]

That distinction matters for lawyers. A search tool that locates a public web page is one thing. A model that can reproduce memorized personal writing from its training corpus is another. The former directs attention outward; the latter demonstrates that information once ingested may later become output. For legal users, the question is not whether the family reunion was emotionally legitimate. The question is whether the same mechanism can make client information, employee details, medical facts, addresses, settlement posture, or privileged strategy less contained than the lawyer assumed.

Split editorial image contrasting a warm family reunion with cold data exposure and hidden personal information

The contrast with Anthropic’s Claude is useful precisely because it is unsentimental. When asked to perform the same search, Claude refused and said, “we don’t want to infringe on anybody’s privacy.”[1] That does not make one vendor safe and another unsafe in any global sense. It does show that model behavior at this boundary is a product and governance choice, not a law of nature. A system can be designed to decline a request that looks like locating a private person. Another can answer.

The Mechanism Is Memorization, Not Magic

The family story should not be reduced to “ChatGPT is good at finding people.” The legally relevant claim is narrower and more uncomfortable: ChatGPT produced exact or near-exact personal text because that text had appeared in material used to train the model. Once the event is stated that way, it stops being a novelty story and starts looking like a live example of training-data memorization.

Diagram of data sources entering an AI model and memorized text being reproduced in response to a user prompt

That mechanism has been documented outside the Singh-Dhamu story. In 2023, Nicholas Carlini and co-authors described a word-repeat attack against ChatGPT in which prompting the model to repeat a word could cause it to emit training data. The researchers reported that more than 5% of ChatGPT outputs in the tested setting were verbatim training data.[2] That is not a claim about every model, every prompt, or every contemporary deployment. It is evidence that, under particular attack conditions, the system could be pushed from ordinary generation into extraction.

The New York Times later reported on research showing that investigators extracted more than 30 employee email addresses from GPT-3.5 Turbo, with 80% accuracy for work addresses.[3] Again, the careful reading is the important one. This was not proof that any user can reliably obtain any person’s contact details on demand. It was proof that personal data can be memorized and elicited from a model that many professionals treated as a general-purpose assistant.

The boundary matters. The extraction studies cited here concern earlier GPT-3.5 Turbo and GPT-4 systems. The sourced materials do not independently verify the same vulnerability profile for newer models such as GPT-4o. A responsible risk analysis should not pretend otherwise. But model updates do not erase the professional-duty problem. Lawyers do not need certainty that the newest model will expose a specific secret before they must ask whether the workflow creates an avoidable confidentiality risk.

A family blog post and a client chronology are not the same kind of document, but the exposure logic rhymes. Lawyers often use generative AI for tasks that feel administratively harmless: summarizing interview notes, cleaning up a demand letter, drafting interrogatories, comparing versions of a settlement proposal, turning a medical timeline into a narrative. Those inputs may contain names, addresses, diagnoses, employment histories, immigration facts, trade secrets, allegations, negotiation positions, and privileged mental impressions.

The risk is not limited to a model blurting out a client’s secret to a stranger tomorrow. It includes several narrower but still serious failures: confidential material being used for vendor training; prompts and outputs being retained longer than the lawyer expects; vendor personnel or subcontractors having access under terms the lawyer has not reviewed; client-identifying facts being inserted into a system without informed consent; and a later privilege fight in which an adversary argues that the lawyer disclosed protected material to a third party without adequate safeguards.

ABA Formal Opinion 512, issued in July 2024, puts the professional-responsibility threshold in concrete terms. It says lawyers must obtain informed consent before inputting client information into a generative AI tool unless the lawyer has a reasonable basis to conclude that the tool’s terms of use, privacy policy, and related safeguards adequately protect the information, including through contractual data protections.[4] That is a practical standard, not an abstract anti-AI warning. The lawyer’s duty turns on what data is being supplied, what the vendor can do with it, how it is retained, who can access it, and whether the client has agreed after being adequately informed.

Legal taskWhat may be exposedRisk question before using a public AI tool
Summarizing intake notesNames, addresses, injuries, family facts, chronologyCan this be done without client-identifying detail, or has the client consented to the vendor workflow?
Drafting a settlement letterLiability theory, demand range, negotiation postureDo the vendor terms bar training and limit retention, access, and onward disclosure?
Reviewing medical recordsHealth information and damages evidenceIs the tool approved for this class of data under firm or client policy?
Preparing deposition questionsAttorney impressions, witness strategy, impeachment materialWould disclosure to this system be defensible if privilege is later challenged?

The privilege problem is especially easy to understate because it often sounds theoretical until it is litigated. Courts may treat waiver questions differently depending on jurisdiction, facts, protective measures, and the nature of the disclosure. ABA guidance does not say that every AI prompt automatically waives privilege. It does make it hard to defend casual prompting when a lawyer has not read the terms, has not disabled training where possible, has not used an enterprise environment with contractual protections, and has not obtained informed consent for client data.

User Chats Are Part Of The Operational Risk

Training-data memorization is only one side of the issue. The other is what happens to the lawyer’s own chats. A Stanford HAI study by King and co-authors, reported in September 2025, found that all six frontier AI developers studied train on user chat data by default, with some retaining data indefinitely.[5] That finding is about vendor practices, not about whether any particular prompt will later be extractable. For law firms and legal departments, however, default training and retention are not administrative trivia. They are the difference between a tool that can be used inside a controlled workflow and a tool that should not receive confidential matter facts at all.

A defensible legal workflow usually begins before the prompt box. Someone must decide which AI environment is approved, whether training on user inputs is disabled by contract or configuration, how long prompts and outputs are stored, whether the vendor can review content for abuse monitoring, whether data is processed outside the expected jurisdiction, and whether the client’s outside-counsel guidelines impose stricter rules. If the answer is “we use the free version because it is convenient,” the analysis has already gone badly.

Two reported episodes show why this is not merely a law-review concern. In an OpenAI community forum post, a user reported that ChatGPT appeared to know a parent’s address without being asked for it.[6] The post is a user report, not a verified judicial finding, so it should be treated cautiously. Still, it is the kind of incident report risk managers collect because it describes the thing users most fear: personal information surfacing outside the context in which the person expected it to remain.

In Australia, a government contractor uploaded personal data relating to up to 3,000 flood victims into ChatGPT, according to ABC reporting.[7] That episode is not the same as proving downstream extraction by another user. Its significance is simpler. People with legitimate work to do will paste sensitive records into convenient tools unless the organization gives them a safer path and a clear prohibition. Lawyers should recognize the pattern. “Do not paste client data into public tools” is obvious advice; obvious advice is often the first control to fail.

What Lawyers Should Treat As Non-Negotiable

The practical response is not to ban every use of generative AI. It is to stop treating prompts as disposable scraps. A client fact pattern entered into a model is a disclosure event unless the lawyer can explain why it is not: the data has been anonymized sufficiently, the tool is covered by appropriate contractual protections, training on inputs is disabled, retention is limited, access is controlled, and the client has consented where consent is required.

  • Use public AI tools only with non-confidential, non-identifying material unless an approved policy says otherwise.
  • Review vendor terms for training, retention, human review, subprocessors, security controls, and deletion rights before client data is entered.
  • Obtain informed client consent when the workflow requires client information and contractual protections are not sufficient under the applicable professional rules.
  • Prefer enterprise or closed legal AI environments with negotiated data protections over consumer accounts used by individual lawyers.
  • Document the approved workflow so privilege and confidentiality decisions are explainable later, not reconstructed after a dispute.

Anonymization also deserves less confidence than it usually receives. Removing a client’s name may not be enough if the prompt preserves a rare job title, a distinctive injury, a small-town address pattern, a filing date, or a unique transaction. The Singh-Dhamu story is a reminder that identity can emerge from surrounding facts. Lawyers should assume that unusual combinations of details may identify a person even when the obvious label is gone.

There is one enforcement point worth narrowing rather than dramatizing. OpenAI’s Italian GDPR matter should not be cited as an active €15 million penalty after a Rome court scrapped the fine in March 2026. The reversal does not make privacy concerns disappear; it simply means the penalty should be described accurately. For legal risk work, accuracy about the enforcement posture is part of the discipline.

The Reunion Proves Enough

The Singh-Dhamu reunion does not prove that today’s newest models will expose any particular client secret. It does not prove that every public chatbot behaves the same way, or that every use of generative AI in legal practice is reckless. It proves something narrower and more relevant: training-data memorization is not a hypothetical mechanism invented by anxious lawyers. It has produced a real human outcome, and independent research has shown related extraction behavior in earlier systems.

That is enough to change the legal analysis. A lawyer who inputs client information into a generative AI tool without consent, contractual protections, and a defensible workflow is not merely experimenting with software. The lawyer is making a confidentiality decision. The fact that the same mechanism once helped reunite a family does not make it less capable of exposing information that another person had reason to keep private.

References

  1. ChatGPT reunited a man with his long-lost half-sister. He had avoided DNA tests and Facebook over privacy fears — The Guardian, July 25, 2026
  2. Extracting Training Data from ChatGPT — not-just-memorization.github.io
  3. How to Make ChatGPT Steal and Tell Your Personal Information — The New York Times, December 22, 2023
  4. ABA issues first ethics guidance on a lawyer’s use of AI tools — American Bar Association, July 2024
  5. AI chatbot privacy concerns, risks research — Stanford News, October 2025
  6. ChatGPT knows my parents address - violation of privacy — OpenAI Community
  7. Can private information uploaded to ChatGPT be found by others? — ABC News, October 8, 2025

Report a correction or tip

Spotted an outdated figure, a misstated fact, or a ruling this case record should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.

Report a correction or tip for this record →