Which Legal Work Can You Safely Delegate to GPT-5.5?
Benchmarks for GPT-5.5 diverge sharply by task: it leads on drafting, contract revision, and document analysis, while compliance review regressed and strict-graded research recall still fails. Legal teams wondering how GPT-5.5 could be used in legal work get a task-by-task deploy-or-verify verdict, grounded in the 2026 benchmark evidence and documented sanction cases.
- Tool
- GPT-5.5
- Benchmark source
- LegalOn, Vals AI, HAQQ, Stanford RegLab, Harvey, Clio, Suprmind/AA-Omniscience
- Hallucination rate
- 3% fabricated citations (HAQQ); 86% refusal-failure hallucination (AA-Omniscience)
- Test methodology
- Mix of human-annotated contract benchmarks, legal-expert evaluations, strict all-pass research grading, and retrieval-benchmark queries
- Test date
- Apr 1, 2026
The safe answer to how GPT-5.5 could be used in legal work starts with a split. Drafting, contract revision, structured summaries, and document analysis are plausible delegation candidates when the human reviewer can check the output against a source file or a known instruction set. Compliance review, open-ended legal research, citation recall, and client-facing advice without verification are not.
That split is not theoretical. In LegalOn’s May 2026 human-annotated contract benchmark, GPT-5.5 improved on contract revision, rising from 85.0% to 87.5% on a 200-item sample, while contract review fell from 80.2% to 77.1% on a 494-item sample. The regression was driven by a 34% increase in false positives, concentrated in business associate agreements and master services agreements.[1] A model that gets better at proposing revisions can still get worse at deciding whether a contract complies with a rule set. For lawyers, that distinction is the whole risk memo.

Other 2026 evaluations support the useful half of that picture. Harvey reported that GPT-5.5 reached 91.7% on its BigLaw Bench, compared with GPT-5.4 at 91.0%, with 43% perfect scores and no task below 0.50.[2] Clio’s legal-expert evaluation placed GPT-5.5 at 87.2%, the top score among the frontier models it tested, with about a 20% relative gain on controlling-authority citation and about a 7% gain on document analysis.[3] Those are vendor evaluations, so they should not be treated as neutral procurement answers. They are still useful signals when read for task shape rather than for marketing temperature.
The cautionary half is just as clear. Vals AI’s Legal Research Bench uses 413 expert-authored questions with strict all-pass grading; GPT-5.5 was reported at approximately 40.38%, while the leading model, Claude Opus 5, was reported at 55.29%. Across models, reconciling conflicting authority was the most reliable failure mode, with only 20.7% all-pass performance.[4] HAQQ’s 300-task, 3,000-answer benchmark gave GPT-5.5 the highest accuracy score among 10 models, 8.41 out of 10, and the lowest hallucinated-citation rate, 3%, but still ranked it fifth overall; across all answers, 24% cited or applied law that did not support the claim.[5] AA-Omniscience, as compiled by Suprmind, reported GPT-5.5 at 57% accuracy with an 86% hallucination rate, where “hallucination” meant confident wrong answers when the model should have refused.[6]
Those percentages are not interchangeable. A fabricated-citation rate, a strict all-pass legal research score, and a refusal-failure hallucination rate measure different dangers. But they point in the same practical direction: GPT-5.5 is much more attractive when the task is bounded and verifiable than when the lawyer is asking it to know, reconcile, or certify the law.
A task-gated verdict
| Legal task | GPT-5.5 use posture | Why |
|---|---|---|
| First drafts of routine legal documents | Deploy with human drafting control | Strong general drafting signals, but final language must be lawyer-owned. |
| Contract revision against supplied instructions | Deploy for proposed edits; verify line by line | LegalOn showed improvement on revision, and the reviewer can compare each change against the contract and playbook. |
| Document analysis and structured summaries | Deploy for extraction and issue spotting; verify against source files | Clio reported gains on document analysis, and the task can be checked against the underlying record. |
| Contract compliance review | Do not delegate final pass | LegalOn found regression and more false positives in review, especially in BAA and MSA contract types. |
| Open-ended legal research | Use only as an assistant to retrieve, organize, or test search paths | Strict research benchmarks remain weak, especially where conflicting authority must be reconciled. |
| Citations, controlling authority, and client-facing conclusions | No unverified output | Sanction cases show that citation failures and unsupported legal claims create professional exposure. |
This table is deliberately less exciting than the benchmark headlines. That is the point. A law firm does not deploy a model score; it deploys a workflow. The question is who checks the work, what they check it against, and whether a bad output creates a cleanup problem before anyone notices.
Where GPT-5.5 is most useful: bounded production work
Drafting
Drafting is the easiest place to justify GPT-5.5 because the output is not supposed to be a legal memory test. A lawyer can give the model a form, a term sheet, a pleading outline, a client chronology, or a clause bank and ask for a first pass. The value is not that the draft is “right.” The value is that the blank page disappears and the human reviewer has something to reshape.
The Harvey and Clio results make GPT-5.5 hard to ignore for this class of work. Harvey’s BigLaw Bench score of 91.7% and Clio’s 87.2% legal-expert score both suggest that, as of August 2026, GPT-5.5 is among the strongest general-purpose models tested on legal production tasks.[2][3] The safer reading is not “let the model draft filings unsupervised.” It is “let the model generate reviewable language where a lawyer already knows the objective, the audience, and the governing constraints.”
Good drafting assignments are narrow. “Draft a short mutual NDA using this approved form and preserve the governing-law clause” is a very different instruction from “draft a motion to dismiss under current federal law.” The first can be compared to a source form and a checklist. The second depends on current law, litigation strategy, record fit, and authority selection. GPT-5.5 may help assemble pieces of the second task, but it should not own it.
Contract revision
Contract revision is the strongest practical case because the reviewer can see the source text, the proposed change, and the instruction that supposedly justified the change. LegalOn’s benchmark is especially useful here because it separates revision from review. GPT-5.5 improved to 87.5% on contract revision in the tested low-reasoning setting, a 2.5 percentage-point gain over the prior result.[1]

That does not make contract revision automatic. It makes it reviewable. The right workflow is to ask GPT-5.5 for proposed redlines against a defined playbook, require an explanation tied to the clause, and have a lawyer or trained contract professional accept, reject, or rewrite each change. If the model touches indemnity, limitation of liability, data processing, termination rights, audit rights, or regulatory language, the reviewer should treat the output as a suggestion rather than a conclusion.
This is also where purpose-built and general-purpose comparisons matter. For a deeper contract-specific discussion, see AI Contract Review Accuracy: What the 2026 Benchmarks Actually Show About Purpose-Built vs. General-Purpose Models. The GPT-5.5-specific point is narrower: revision tasks can be delegated earlier than compliance judgments because the reviewer can trace the model’s work back to a clause and an instruction.
Document analysis and structured summaries
Document analysis sits between drafting and legal research. It is useful when the assignment is to extract, summarize, classify, or compare information contained in a defined set of documents. Clio reported about a 7% relative gain for GPT-5.5 on document analysis, which is consistent with the model’s broader strength on bounded legal work.[3]
The safest uses are source-grounded: summarize this deposition transcript; extract all termination provisions from these agreements; build a chronology from this production set; compare this revised clause to the prior version; identify obligations assigned to the vendor in this statement of work. These tasks still require verification, but the verification is visible. The reviewer can search the source, check quotations, and inspect whether the summary left out something material.
A bad document summary can still cause harm. It can omit a qualifier, flatten a disputed fact, or convert a witness’s uncertainty into a clean admission. But the failure mode is usually discoverable if the workflow requires pin cites, source excerpts, and spot-checking by someone who understands the matter. GPT-5.5 should be asked to show its work in a format that makes review faster, not in a format that makes review feel unnecessary.
Where the evidence says to stop
Compliance review
Compliance review is where GPT-5.5’s fluent legal style becomes more dangerous. The LegalOn result is the cleanest warning: the same model that improved on contract revision regressed on contract review, falling to 77.1% and producing more false positives.[1] A false positive in this setting is not a harmless extra sentence. It can tell a lawyer or business team that a clause has a problem when it does not, or that a contract fails a requirement when the issue is really interpretive, contextual, or absent.
Compliance work often requires the model to map legal requirements, company policy, contract type, regulatory definitions, exceptions, and risk tolerance onto messy text. That is not just pattern matching. It is judgment under consequences. GPT-5.5 can assist by locating relevant clauses, producing checklists, or drafting a memo template, but the final compliance call should remain with a qualified reviewer.
Open-ended research and authority reconciliation
Open-ended legal research is the other no-delegation zone. The Vals AI result is severe because of its grading method: strict all-pass scoring gives no credit for an answer that is partly useful but legally incomplete. On that measure, GPT-5.5’s reported score of approximately 40.38% is not enough for unsupervised research, and the pooled 20.7% all-pass performance on reconciling conflicting authority shows exactly where the danger lies.[4]
A model can sound most lawyerly when it is least entitled to be trusted. Reconciling a split, identifying controlling authority, distinguishing dicta, recognizing procedural posture, and knowing when no good answer exists are not cosmetic tasks. They are the work product. GPT-5.5 can help generate search terms, compare cases supplied by the lawyer, or turn verified research into a readable outline. It should not be treated as the source of legal authority.
Even retrieval-grounded legal research systems do not eliminate the problem. Stanford RegLab and HAI found that tools marketed around reduced hallucination still failed on more than 17% of benchmark queries, with one system failing above 34%.[7] That evidence matters because retrieval is supposed to be the safety layer. If legal-specific research tools still miss that often, a general-purpose model should not be granted a lower standard merely because it produces a polished answer.
Why buyers are asking now
GPT-5.5 arrived in April 2026 with a product frame that made legal teams pay attention: Thinking, Pro, and Instant tiers; a 1 million-token API context window; and published API pricing of $5 per 1 million input tokens and $30 per 1 million output tokens for one tier, and $30 per 1 million input tokens and $180 per 1 million output tokens for another.[8] Those facts matter because they make large document sets and longer legal workflows feel operationally possible. They do not answer whether any particular workflow is safe.
The timing also matters. GPT-5.6 shipped on July 9, 2026, so the GPT-5.5 benchmark record is a dated snapshot, not a current-model ranking. It is still useful for deciding how GPT-5.5 itself should be deployed where firms have already tested, licensed, or built around it.
A long context window is useful for loading more contracts, transcripts, policies, or diligence materials. It also lets teams create larger mistakes. More context does not guarantee that the model will identify the legally decisive passage, reconcile inconsistent instructions, or refuse a question it cannot answer. Procurement teams should evaluate context length as capacity, not as reliability.
The pricing question belongs in the same lane. Lower unit cost can justify more experimentation, more internal pilots, and more workflow redesign. It cannot justify removing review from tasks where the benchmark evidence shows persistent failure. For a broader buyer-risk discussion, see The AI Selloff Repriced Legal AI Vendors, Not Risk.
The failure modes map directly to professional duty
The legal profession already has a live record of what happens when lawyers treat generated legal text as if it verified itself. The sanction history runs from Mata v. Avianca, where lawyers were sanctioned $5,000 in 2023, to Couvrette v. Wisnovsky, with sanctions reported around $109,700, to Withers v. City of Aberdeen, where a June 8, 2026 order canceled a trial and suspended two lead attorneys for two years. Other listed incidents include Wadsworth v. Walmart, where an in-house platform fabricated 8 of 9 citations, and Barteca v. Tacobarn, with a $3,500 sanction plus bar referral.[9]
The lesson is not that every use of GPT-5.5 is reckless. The lesson is that the lawyer remains the person holding the filing, the opinion, the contract, or the advice. ABA Formal Opinion 512, issued July 29, 2024, frames generative AI use through existing duties including competence, confidentiality, communication, supervision, and fees.[10] For more on that duties framework, see ABA Formal Opinion 512: What Generative AI Ethics Rules Actually Require of Attorneys and What the ABA and State Bars Require of Lawyers Using AI.
This is where benchmark interpretation becomes a legal-risk exercise. A 3% fabricated-citation rate in one benchmark is impressive compared with worse systems, but it is still intolerable if the workflow lets one fabricated citation reach a court. An 87.5% contract-revision score can support a supervised redline workflow, but it does not support signing off on the model’s edits without review. A 1 million-token context window can reduce copy-and-paste friction, but it does not make a model competent to decide unresolved law.
For confidentiality and privilege, the question is not just whether GPT-5.5 can perform the task. It is where the data goes, who can access it, whether inputs are retained or used for training, and whether the lawyer has client consent where required. For more on that layer, see The Ethics of Free AI for Lawyers. The model’s legal benchmark score does not answer those deployment questions.
A workable GPT-5.5 legal workflow
The safest GPT-5.5 rollout is not organized around enthusiasm or fear. It is organized around task gates.
- Allow GPT-5.5 to draft from approved templates, rewrite for tone, summarize bounded source materials, extract defined fields, and propose contract revisions against a supplied playbook.
- Require source-linked review for every output that depends on a document set, contract, transcript, policy, regulation, or case.
- Block unsupervised use for legal research conclusions, controlling-authority statements, citation generation, compliance determinations, settlement advice, and client-facing legal answers.
- Log user instructions, source materials, model outputs, reviewer decisions, and final edits for any workflow likely to affect client advice or filed work.
- Test the model on the firm’s own documents and matter types before procurement claims are treated as deployment evidence.
The review standard should be stricter when the model’s output contains legal propositions rather than document observations. “Section 9.2 says either party may terminate on 30 days’ notice” can be checked against Section 9.2. “This termination right is enforceable under California law” requires legal analysis. “The vendor’s indemnity obligation excludes third-party IP claims” can be checked against the indemnity clause. “The agreement complies with HIPAA business associate requirements” crosses into compliance judgment.
This line also helps assign work inside the legal team. Junior lawyers, contract managers, and knowledge staff can use GPT-5.5 to accelerate production when there is a review path and a known source of truth. Senior lawyers should own tasks where the model’s answer would otherwise substitute for legal judgment. For contract-analysis-specific professional responsibility issues, see The Professional Responsibility Guide to AI Contract Analysis.
The narrow answer
As of August 2026, GPT-5.5 can be used in legal work for bounded, checkable production tasks: first drafts, contract redlines, structured summaries, extraction, comparison, and document analysis. It should be treated as a drafting and review accelerator, not as a legal authority engine.
The no-delegation zone remains equally clear: precise legal recall, reconciliation of conflicting authority, compliance determinations, citation reliability, and unverified client-facing conclusions. Those tasks stay with humans because the professional duty to verify does not move to the model. This article is informational and is not legal advice.
References
- GPT-5.5 Takes a Step Forward, But the Picture Is More Complicated Than It Looks, LegalOn, May 5 2026
- GPT-5.5 Research Preview Results, Harvey, April 23 2026
- GPT-5.5 for Agentic Legal Work, Clio, updated May 28 2026
- Legal Research Bench, Vals AI
- Best AI for Legal Work Benchmark, HAQQ, June 5 2026
- AI Hallucination Rates and Benchmarks, Suprmind
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Stanford RegLab
- Introducing GPT-5.5, OpenAI, April 23 2026
- Damien Charlotin: AI Hallucinations Cases, Damien Charlotin
- ABA issues first ethics guidance on a lawyer’s use of AI tools, American Bar Association, July 29 2024
Chronological incident history
No sanction cases have named this tool in the tracked record set to date. This does not imply the tool is safe — see Risk Digest for ongoing monitoring.
← Compare peer toolsReport a correction or tip
Spotted an outdated figure, a misstated fact, or a ruling this tool profile should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.
Report a correction or tip for this record →