Skip to content

Workflows

Custom Gemini Gems Don't Reduce Legal AI Risk

Lawyers may assume custom Gemini Gems are safer for legal workflow tasks because instructions are tailored, but the Gem inherits the same hallucination, confidentiality, and supervision risks as the underlying Gemini tier. This article examines the research and documented sanction cases to show why customization is not a risk mitigation measure.

By Editorial TeamUpdated Jul 25, 2026
Applicable role
attorney
Workflow stage
review
Primary source
ABA Formal Opinion 512

A custom Gem can make a legal AI workflow look disciplined. The research prompt is already written. The drafting conventions are embedded. The internal playbook sits beside the model instead of in someone’s saved notes. For a litigation team trying to standardize how associates summarize cases or how paralegals triage discovery questions, the appeal is obvious.

That is also where the false comfort begins. Google describes Gems as customized versions of Gemini that can be configured with instructions for recurring tasks; its support materials explain how users can create and manage those Gems inside Gemini. Those features may improve consistency, but they do not change what courts and ethics authorities usually examine after an AI error: whether the filing was verified, whether client information was protected, and which lawyer supervised the work. [1][2]

This verification-workflows analysis is current as of July 25, 2026, and is not legal advice. The narrower point is operational: a custom Gemini Gem used as a legal workflow advisor is still Gemini running under a particular account tier, with a human lawyer remaining responsible for the use made of its output.

Cracked translucent gem on a law library desk with legal books, documents, a gavel, a padlock, and a sanction stamp visible inside

What customization changes, and what it leaves untouched

Custom instructions are useful. They can tell the Gem to use a firm’s preferred format, ask clarifying questions before drafting, avoid certain language, route the user through a checklist, or remind the user to attach source material. A knowledge-management lawyer might reasonably use a Gem to reduce variation in first-pass work.

But those instructions are not a liability shield. They do not make a generated citation real. They do not rewrite the data-retention terms that apply to the account. They do not turn the model into a lawyer, paralegal, vendor, or employee who can bear professional responsibility. They sit on top of the underlying service.

Risk questionWhat a custom Gem may improveWhat it does not eliminate
Citation reliabilityPrompt consistency, source-requesting habits, formatting disciplineThe need to verify every case, quotation, statute, docket reference, and procedural statement
ConfidentialityInternal reminders not to enter client data, routing to approved workflowsThe privacy terms, retention rules, review practices, and contractual controls of the actual Gemini tier
SupervisionRepeatable instructions for associates, staff, and legal operations usersThe supervising lawyer’s responsibility for AI-assisted work product

That distinction matters because the reported sanction cases do not turn on whether the lawyer used a clean interface, a clever prompt, or a custom wrapper. They turn on what reached the court and what the lawyer did before filing it.

Citation hallucination remains the first hard stop

The strongest argument against treating a custom Gem as a safety control is the hallucination record. Stanford RegLab’s benchmark work found hallucination rates of 17% to 34% even in paid legal research tools. That finding is especially important because those tools are closer to legal research products than a general chatbot is; a customized Gem should not be assumed to outperform the underlying model merely because the prompt sounds lawyerly. [3]

The public case record is no longer anecdotal. The HAQQ/Charlotin AI hallucination database documented 1,598 hallucination-related cases globally through June 2026, with more than 1,148 in the United States alone. Those are documented floors, not a complete census of every AI-assisted legal error. [4]

Norton Rose Fulbright’s 2026 AI sanctions survey and EDRM’s Q1 2026 penalty reporting describe the same judicial pattern from different angles: courts are imposing real consequences for fabricated authorities, misleading explanations, and failed verification. EDRM reported more than $145,000 in Q1 2026 AI-related penalties alone. [5][6]

Editorial diagram showing three risk pathways from a gem icon to broken citations, exposed data, and a liability stamp

The sanction record is about filings and verification

The named cases are useful because they show what courts punish. In Fletcher v. Experian, the sanction was $2,500 and the court’s concern included misleading the court about AI use. In Whiting v. City of Athens, sanctions reached $15,000 per attorney plus fees. In Farris, the lawyer was removed from the case and fees were denied despite candor about the AI issue. Mostafavi involved a $10,000 sanction tied to 21 fabricated quotes. An Oregon Court of Appeals fee schedule imposed $500 per fabricated citation and $1,000 per fabricated quotation. [4][5][6]

Those details should cure any casual confidence that the problem is just embarrassing cleanup work. A fabricated quotation can become a fee problem. A fake citation can become a sanctions problem. A misleading explanation about how the error happened can become its own aggravating fact.

The current record does not show a court sanctioning a lawyer solely because the lawyer used a custom Gemini Gem. That absence matters. The risk analysis here is not based on a direct custom-Gem sanctions ruling; it is based on the documented treatment of AI-generated legal work when false material reaches a court. Public trackers and surveys include cases involving AI tools, including Gemini among named tools, but they do not create a separate doctrine for customized wrappers. [4][5]

That is the practical point. If a Gem fabricates a case name because the user asked it to draft a motion section, the court will not need to decide whether the Gem was well designed. The first questions will be more ordinary: Who signed the filing? Who checked the authorities? Who represented the result to the tribunal? Who had a chance to catch the error before it entered the record?

Supervision does not move inside the Gem

ABA Formal Opinion 512, issued in July 2024, treats the lawyer’s use of generative AI as part of existing professional duties rather than as a separate technology exception. Its supervision analysis places AI tools within the lawyer’s responsibility to supervise nonlawyer assistance under Rule 5.3. The tool may generate text, but the lawyer remains responsible for how that text is used. [7]

A custom Gem does not become a legally accountable assistant because it has a firm-approved prompt. It cannot certify that authorities were checked. It cannot exercise professional judgment about whether a passage overstates a holding. It cannot decide whether a client’s confidential facts should be used in a particular system. The human chain of responsibility remains visible even when the user interface hides the model behind a tailored name.

The Singapore High Court’s decision in Tan Hai Peng v Tan Cheong Joo is a useful warning outside the U.S. ethics-opinion frame: a supervising partner who signs unchecked AI-assisted work can be personally liable. The point travels well even if local doctrine varies, because it rests on a familiar procedural expectation. A signature is not a routing stamp. [8]

There is also an uncomfortable institutional asymmetry. A 2026 Northwestern study found that 45.5% of federal judges reported that their court administration had not provided AI training, yet courts are still sanctioning attorneys for AI-related errors. Lawyers should not assume that a court’s limited institutional AI training will translate into indulgence for unverified AI output. [9]

For a supervising lawyer asked to approve a Gem after a practice group has already started using it, the useful approval question is not, “Are the instructions good?” The better question is, “What human review is mandatory before any output leaves the firm or affects a client matter?” If that review cannot be described, assigned, and documented, the customization has solved the wrong problem.

Confidentiality depends on the tier, not the label

Confidentiality is where product language can become especially misleading. A Gem named “Litigation Research Assistant” may feel internal, but the privacy analysis starts with the account tier and applicable contract. Consumer Gemini and Workspace or enterprise deployments are not the same confidentiality environment.

For consumer Gemini accounts, the research materials identify prompt and response retention for up to 36 months with possible human review. Google describes data as de-identified in this context, but de-identified is not the same thing as anonymous, and it is not a substitute for a privilege analysis before entering client information. [10]

ABA Formal Opinion 512 requires lawyers to fully consider confidentiality before entering client information into any generative AI tool. That obligation is not reduced because the lawyer enters the information through a custom Gem rather than a blank Gemini window. [7]

Independent legal-technology privacy analyses reach the same practical conclusion from a procurement angle: general-purpose chatbot tiers do not provide the structural data isolation that legal privilege requires unless the firm has verified enterprise contractual protections. Spellbook’s Gemini privacy materials and JLE’s privacy analysis both emphasize tier-specific review rather than trust in the chatbot brand. [11][12]

Google’s Gemini Enterprise materials describe legal research use cases and enterprise controls, including contractual data isolation, client-side encryption, and audit logs. Those features are materially different from consumer use, but they still require verification at the firm level: whether the correct add-on is activated, whether the Data Processing Agreement covers the intended AI use case, whether audit logs are available to the right administrators, and whether the specific matter information is permitted in that environment. [13]

The safe formulation is narrow. Enterprise controls can support an approved legal AI workflow when the contract, settings, retention rules, and use case have been checked. The existence of a custom Gem, by itself, says none of that.

Where a custom Gem can still be useful

None of this makes custom Gems useless. A firm may sensibly use them for repeatable, low-risk workflow structure: converting deposition notes into a standard internal format, reminding users to supply source documents, generating issue-spotting questions for human review, or routing a draft through required verification steps before it can be used externally.

The Advocate Magazine’s priming-executing-verifying workflow is helpful as a practical guardrail: frame the task, run the AI-assisted work, then verify the result before relying on it. It is not an official bar standard, and it should not be treated as one. Its value is that it keeps verification as a distinct human step rather than letting a polished prompt masquerade as review. [14]

A firm-approved Gem should therefore carry operational limits in the same place it carries prompt instructions. For citation work, the user must check every authority in an authoritative legal database or court source before filing or advising. For summarization, the user must compare the summary against the source document rather than relying on the model’s confidence. For drafting, the signer must review the legal standard, factual record, quotations, and procedural posture. For confidential information, the user must know which Gemini tier is being used and whether that tier is approved for the matter data involved.

The documentation should be as plain as the risk. Record the approved account tier, the permitted use cases, the prohibited inputs, the verification steps, and the person responsible for final review. If a court, client, insurer, or ethics authority later asks what happened, a firm will not want to answer with a screenshot of a well-written Gem instruction.

Treat a custom Gemini Gem exactly like the underlying Gemini tier for verification, confidentiality, and supervision. The wrapper may improve workflow discipline; it does not reduce the lawyer’s duty to check the work, protect the client’s information, and stand behind what leaves the office.

References

  1. Tips for creating custom Gems — Google.
  2. Create & manage Gems — Google.
  3. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools — Stanford RegLab.
  4. AI Hallucination Cases: The 1,598-Case Sanctions Tracker — HAQQ.
  5. AI in litigation: Update on Gen AI sanctions in 2026 — Norton Rose Fulbright.
  6. The AI Sanction Wave: $145K in Q1 Penalties — EDRM.
  7. ABA Formal Opinion 512 — American Bar Association, July 2024.
  8. Tan Hai Peng v Tan Cheong Joo — Singapore High Court.
  9. Northwestern 2026 study on federal judges and AI training — Northwestern.
  10. Gemini Apps Privacy Notice — Google.
  11. Gemini for Lawyers; Is Gemini Safe for Legal Work? — Spellbook.
  12. Gemini and Privacy for Lawyers — JLE.
  13. Use case: Conduct legal research | Gemini Enterprise — Google.
  14. Priming-executing-verifying workflow — Advocate Magazine.

Grounded in

This procedure is grounded in ABA Formal Opinion 512, independent of any single documented case. See the Regulation tracker for the governing text.

Cases this step would have prevented

No cases have been explicitly linked to this checklist yet. See Risk Digest for documented incidents generally.

← Back to Workflows

Report a correction or tip

Spotted an outdated figure, a misstated fact, or a ruling this workflow checklist should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.

Report a correction or tip for this record →