Skip to content
Lex Machina Review logoLex Machina Review
Menu

Evaluations

GPT-5.6 Leads DeepSeek V4 on Legal Research, Neither Is Safe

The July 2026 legal-research benchmarks show GPT-5.6 leading DeepSeek V4 by a wide margin, yet the best generalist still fails more than half of strict checklist-style tasks. This evaluation turns the scores into a procurement and filing risk judgment, mapping each model family's distinct failure modes and the verification obligations that survive either choice.

Hallucination rate
Not measured / undisclosed

For legal research in Q3 2026, GPT-5.6 is the safer general-purpose model family than DeepSeek V4, and the margin is not close. On the July 23, 2026 Legal Research Bench snapshot attributed to Vals AI and mirrored by BenchLM, GPT-5.6 Sol scored 48.08% all-pass accuracy, while DeepSeek V4 Pro scored 23.08% — roughly a 25-point gap, with Sol ranked 3rd and DeepSeek V4 Pro ranked 20th among 27 models.[1][2] That is a real procurement signal. It is not a filing clearance.

The uncomfortable part is the denominator: even the leading GPT-5.6 variant failed more than half of strict legal-research tasks on that benchmark.[1][2] A litigation team can use that gap to reject a weaker model for high-risk research. It cannot use the same table to let a model’s answer move into a memo, client advice, or court filing without independent verification.

Two AI orbs on a dark track, one far ahead, with a red safety threshold above both and a courtroom scale in the background

One trust caveat belongs near the top. The available Vals figures should be date-stamped and rechecked directly on vals.ai before a procurement memo treats them as current, because the captured table is available through the BenchLM mirror and snippets rather than a clean direct crawl of the JavaScript-rendered Vals page.[1][2] In a fast-moving model market, “GPT-5.6” is too vague; the relevant comparison here is GPT-5.6 Sol, Terra, and Luna against DeepSeek V4 Pro, with the V4 Pro result understood as the named legal-benchmark row rather than a proxy for every DeepSeek deployment mode.

The benchmark gap is large, but the winning score is still failing work

The July Legal Research Bench numbers put three GPT-5.6 variants above DeepSeek V4 Pro by a meaningful margin. Sol leads this comparison; Terra and Luna also clear DeepSeek V4 Pro comfortably.[1][2]

Vals AI Legal Research Bench figures from the July 23, 2026 snapshot, with mirror caveat.
Model variantLegal Research Bench all-pass accuracyRank signal in cited snapshot
GPT-5.6 Sol48.08%3rd of 27
GPT-5.6 Terra40.87%Above DeepSeek V4 Pro
GPT-5.6 Luna36.54%Above DeepSeek V4 Pro
DeepSeek V4 Pro23.08%20th of 27

All-pass accuracy matters because legal research often fails at the weakest required element. A response can identify the right general doctrine, sound fluent, and still be unusable if it cites the wrong authority, ignores a jurisdictional limit, misses a negative-treatment instruction, or answers only part of the prompt. A “mostly right” answer may be helpful during exploration. It is a poor unit of measurement for a task that ends with a lawyer signing a document.

That is why the Sol-versus-DeepSeek gap should not be minimized. A model that passes 48.08% of strict tasks is materially less dangerous than one that passes 23.08% if both are being considered for the same legal-research workflow.[1][2] But the same number also tells the KM lawyer where the procurement memo must stop. Sol is not close to a “research associate replacement” score. It is a better candidate for a controlled workflow: issue spotting, first-pass authority collection, draft research trails, and structured prompts that are later checked against source law.

The adjacent cost question should stay adjacent. If the buyer is comparing token prices, model routing, or DeepSeek V4 Flash rather than V4 Pro, that belongs in the separate pricing and Flash benchmark discussion at DeepSeek V4 Flash vs OpenAI pricing and DeepSeek V4 Flash agent benchmarks. This comparison is narrower: when the task is legal research, the July legal-domain evidence favors GPT-5.6 by a wide measured margin.

A second leaderboard confirms the direction, then exposes the drafting problem

The July 2026 Legal Benchmarks AI practitioner leaderboard points the same way, though with different task construction. Across 63 legal tasks, it reported GPT-5.6 Sol reliability at 44.1% and DeepSeek V4 Pro at 26.5%.[3] That is not the same benchmark as Vals, and the percentages should not be blended into a single average. The useful point is directional consistency: GPT-5.6 Sol remains ahead when the measurement frame changes.

The same leaderboard also prevents the easy version of the story. Its drafting finding is the kind of result that should make supervisors slower, not calmer: GPT-5.6’s problem is often polished-but-wrong output, including drafts where more than half miss at least one instruction.[3] That failure mode is expensive because it survives a skim. The prose looks like work product. The formatting is plausible. The defect appears only when someone compares the output against the prompt, the authorities, and the required jurisdictional or procedural constraints.

Split illustration showing a polished legal document with a hidden error on one side and tangled citation pins with a cracked shield and padlock on the other

That distinction matters in staffing. DeepSeek V4 Pro’s legal-research score raises a front-end selection question: why put the weaker measured model into a workflow where citation accuracy is central? GPT-5.6 Sol raises a back-end review question: who is responsible for finding the one instruction it missed after the draft has already been made easy to like?

DeepSeek V4 Pro’s risk is not just a lower score

DeepSeek’s July legal-benchmark position is already enough to make it hard to justify for unsupervised or lightly supervised legal research. The supporting hallucination evidence adds a citation-specific warning, but it has to be used carefully. Digital Applied’s April 2026 5,000-prompt study is non-peer-reviewed vendor research, and it tested DeepSeek V4 against GPT-5.5 rather than GPT-5.6.[4] It should be treated as a stress signal, not a final scientific ranking.

With that label attached, the results are still relevant to legal work. The study found citation accuracy to be the worst frontier task family, with a 12.4% average hallucination rate; DeepSeek V4 was worst among the five tested models at 19.1% by default and 15.7% with chain-of-thought, while the best reported configuration was GPT-5.5 Pro with extended thinking at 6.8%.[4] The study also reported that retrieval grounding cut citation hallucinations by 75–90%, while prompting alone reduced them by only 5–15%.[4]

For legal operations, that last comparison is more useful than another debate about prompt cleverness. If citation reliability is the bottleneck, the control is not a longer instruction like “do not hallucinate.” The control is grounded retrieval, source display, citation checking, and a human reviewer who opens the cited authority. Prompting may improve behavior at the margin; it does not turn an ungrounded general-purpose model into a filing-safe research system.

NIST CAISI’s May 1, 2026 evaluation gives DeepSeek V4 Pro a different kind of context. CAISI described it as the most capable PRC model it had tested, while estimating it at about eight months behind the U.S. frontier, with IRT Elo 800±28 compared with GPT-5.5 at 1260±28.[5] That is useful frontier-position evidence. It does not decide legal safety, and it does not erase the legal-domain results. A model can be impressive in general capability and still be the wrong choice for citation-heavy legal research.

DeepSeek also tends to carry procurement questions that sit outside the benchmark table: data handling, confidentiality, deployment geography, client restrictions, and firm policy. Those questions are not solved by a legal-research score, whether good or bad. Teams already comparing model governance patterns may want the related procurement frame in Claude Opus 5, Fable 5, and compliance, but the minimum point here is simpler: a weaker legal benchmark plus a harder confidentiality review is not an attractive risk package for research that may flow into advice or filings.

GPT-5.6’s safer profile still breaks at review

GPT-5.6 Sol is the defensible choice if the buyer insists on one of these general-purpose model families for legal research. The word “defensible” is doing work. It means the selection can be supported by named legal-domain benchmarks. It does not mean the model’s legal output can be trusted without verification.

The surviving GPT-5.6 risk is not that every answer looks wild. It is that many bad answers look finished. A clean memo section that missed a limiting instruction creates a different burden from a visibly broken response. The associate or KM reviewer has to reconstruct the prompt, check each required element, open each cited authority, and decide whether the answer is wrong, incomplete, or merely unsupported. The time cost moves downstream, often to the person least responsible for the purchase decision.

That is also why comparisons such as Claude vs. ChatGPT for legal work keep returning to verification rather than model personality. The review layer is not a belt-and-suspenders preference. It is the part of the system that turns model output into legal work product.

A document passing through citation verification, instruction review, and final legal sign-off checkpoints before a human hand stamps it

The sanction record is model-agnostic, and that makes it more relevant

There is no need to pretend the sanction record specifically condemns DeepSeek V4 or GPT-5.6. Damien Charlotin’s AI hallucination cases database, updated July 29, 2026, listed 1,811 documented cases, including 1,252 in the United States, but no sanction decision in the cited material names either DeepSeek V4 or GPT-5.6.[6] The exposure is tool-agnostic: lawyers get in trouble for what they file, certify, deny, or fail to correct.

That is the more durable risk signal anyway. The model name will change; the duty to verify authority will not. For recent sanction patterns and enforcement principles, see the site’s Risk Digest coverage of AI-generated citation sanctions and AI citation hallucination sanctions in federal courts. The practical lesson is not “avoid the named model in the last bad case.” It is “do not file unverified authority because a system made it sound real.”

Legal-specific retrieval systems do not make that problem disappear. Stanford HAI and RegLab reported that RAG-based legal tools still hallucinated: Lexis+ AI and Ask Practical Law AI at more than 17%, and Westlaw AI-Assisted Research at more than 34% on the studied benchmark queries.[7] The same line of work warns that misgrounded citations can be more pernicious than fabricated cases because they attach an answer to a real-looking source trail.[7]

That finding should discipline how teams read generalist-model benchmarks. If dedicated legal RAG tools still produce wrong or misgrounded answers, a general-purpose model with a sub-50% all-pass legal-research score is not a candidate for unchecked citation work. It may be a useful component. It is not the control environment.

Professional-responsibility guidance points in the same direction. ABA Formal Opinion 512 and the ARDC guidance emphasize competent, confidential, supervised use of generative AI in legal practice.[8] A 2025 sanctions review likewise frames Rule 11(b) obligations in tool-agnostic terms: the lawyer’s certification obligation remains, and courts have treated denial or failure to correct as aggravating conduct in hallucination matters.[9]

What a procurement memo can safely say

A careful procurement memo does not need to flatten the evidence into “GPT good, DeepSeek bad.” It can say something narrower and more useful: on the July 2026 legal-domain evidence available, GPT-5.6 Sol is substantially stronger than DeepSeek V4 Pro for legal research; Terra and Luna also beat DeepSeek V4 Pro on the cited Vals snapshot; none of the tested general-purpose options clears a threshold for unsupervised legal-research output.[1][2][3]

The memo should also separate adoption from effectiveness. A firm may adopt a model for brainstorming, summarizing non-confidential public materials, or generating a first research plan. That does not establish that the same model is effective for citation-safe research. A benchmark score is closer to effectiveness evidence, but only within the benchmark’s task design, model variant, date, and settings.

  • For DeepSeek V4 Pro: the July legal-research evidence supports avoiding unsupervised or lightly supervised use for citation-heavy legal research. If used at all, it belongs in low-risk exploratory work with strict data controls and full downstream verification.
  • For GPT-5.6 Sol: the evidence supports preferential selection over DeepSeek V4 Pro for controlled legal-research workflows, not direct filing use. The review plan must assume missed instructions and plausible but wrong citations.
  • For GPT-5.6 Terra and Luna: the Vals snapshot places them above DeepSeek V4 Pro, but below Sol. If procurement is risk-weighted rather than price-weighted, Sol is the variant to evaluate first.
  • For DeepSeek V4 Flash: do not borrow V4 Pro’s legal-benchmark row or infer a legal-research clearance from pricing claims. Treat Flash as a separate evidence question.

The verification budget should be explicit. It belongs next to the license fee, not hidden in associate time after rollout. For teams building that review layer, the relevant operational question is not whether a model can produce a persuasive draft; it is whether the workflow forces citation opening, negative-treatment checks, instruction-by-instruction comparison, privilege and confidentiality review, and final attorney sign-off. The verification-gap budgeting issue is discussed more directly in Dunia AI funding and the verification gap, and the cost of review also appears in Opus 5 API legal task costs.

The risk classification

DeepSeek V4 Pro is hard to justify for unsupervised or lightly supervised legal research on the July 2026 evidence. The Vals all-pass result is low, the Legal Benchmarks AI result is also behind GPT-5.6 Sol, and the citation-hallucination stress evidence points in the wrong direction even after the necessary caveats.[1][2][3][4]

GPT-5.6 Sol is the stronger and more defensible general-purpose choice for legal research, with Terra and Luna still ahead of DeepSeek V4 Pro in the cited Vals snapshot.[1][2] The safety classification remains supervised-use only. Independent citation verification, instruction-by-instruction review, confidentiality controls, and final attorney responsibility survive the model choice.

The procurement answer to “DeepSeek V4 vs GPT-5.6 for legal research” is therefore relative, not absolute: choose GPT-5.6 if forced to choose between these model families for legal-research support, but do not let the benchmark winner become the filing authority.

References

  1. Legal Research Bench, Vals AI, July 23, 2026
  2. Legal Research Bench, BenchLM, July 23, 2026
  3. Leaderboard, Legal Benchmarks AI, July 2026
  4. AI Model Hallucination Rate Benchmarks 2026 Study, Digital Applied, April 2026
  5. CAISI Evaluation of DeepSeek V4 Pro, NIST, May 1, 2026
  6. AI Hallucination Cases Database, Damien Charlotin, updated July 29, 2026
  7. AI on Trial: Legal Models Hallucinate in 1 out of 6 or More Benchmarking Queries, Stanford HAI
  8. ABA Formal Opinion 512 and ARDC guidance, Illinois Courts
  9. AI & IP Year in Review—AI Hallucinations in Court Filings and Orders: A 2025 Review of Sanctions Across the Courts and Rule Proposals, Sterne Kessler, 2025

Chronological incident history

No sanction cases have named this tool in the tracked record set to date. This does not imply the tool is safe — see Risk Digest for ongoing monitoring.

← Compare peer tools

Report a correction or tip

Spotted an outdated figure, a misstated fact, or a ruling this tool profile should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.

Report a correction or tip for this record →
Blogarama - Blog Directory