Skip to content
Lex Machina Review logoLex Machina Review
Menu

Evaluations

DeepSeek V4 Flash agent benchmark scorecard for legal teams

DeepSeek V4 Flash posts class-leading agentic-coding scores at outlier-low prices, yet has no published legal-domain benchmark and shows near-record overconfidence on knowledge tasks. This scorecard separates provider-reported from independent results and shows legal teams exactly where V4 Flash needs a grounding and cite-verification layer before use.

Tool
DeepSeek V4 Flash
Benchmark source
DeepSeek provider tables; BenchLM; Artificial Analysis; CAISI/NIST
Hallucination rate
approx. 96% (AA-Omniscience)
Test methodology
Mixed benchmark ledger separating provider-reported rows from independent/adjacent evidence, with thinking-mode labels and sibling-model inferences flagged; no Flash-specific legal benchmark.
Test date
Jul 31, 2026

DeepSeek V4 Flash is the kind of model that makes an engineering team pause: an open-weight MoE model, a 1M-token context window, MIT licensing, and API pricing listed at $0.14 input and $0.28 output per 1M tokens. It also posts very strong provider-reported agentic-coding results. The legal problem is that the evidence gets thin exactly where a legal team needs it to get specific: legal authority, citation reliability, and calibrated refusal. As of this record date, the public record contains no Flash-specific legal-domain benchmark row, while available knowledge-calibration evidence points to a model that may answer far too often when it should abstain. [1][2][3][4][5]

Split editorial image showing fast agentic coding strength on one side and uncertain legal verification risk on the other

The immediate procurement answer is therefore narrow. V4 Flash is worth testing as a low-cost coding, agent, and workflow component. It should not be treated as a legal authority unless the legal authority is supplied by a retrieval or grounding layer and the citations are independently verified. Any legal-risk estimate for Flash is inferred from sibling-model and adjacent evidence, not measured directly on a published Flash legal benchmark.

Model record

Snapshot date for this dossier: July 31, 2026.
FieldRecord for DeepSeek-V4-FlashEvidence label
ReleaseApril 24, 2026 [1]Provider release note
Architecture footprint284B total parameters / 13B active MoE [2]BenchLM model profile
Context length1M tokens [1][2]Provider release note and model profile
LicenseMIT [2]BenchLM model profile
API price$0.14 input / $0.28 output per 1M tokens [1]Provider release note
Procurement status for legal workTestable as an assisted workhorse; not acceptable as a stand-alone legal authority without grounding and cite verificationEvaluation judgment from the benchmark record

That price-performance profile is not a side detail. At legal-department scale, low inference cost changes which experiments are affordable: bulk document routing, coding-heavy integrations, contract-data extraction pipelines, and agentic testing harnesses all become easier to justify. The mistake would be to let that operational attractiveness turn every benchmark row into a legal-reliability row.

Benchmark ledger: what the scores actually measure

A single benchmark table is not enough for this model. V4 Flash’s profile changes by task type and, in several rows, by reasoning mode. A legal buyer needs the label before the score: provider-reported or independent, coding or knowledge, closed-book or grounded, non-think or Think Max.

Three-part scorecard illustration showing agentic coding strength, knowledge overconfidence, and a missing legal benchmark slot
The ledger separates provider-reported model tables from independent or adjacent evidence. It does not convert non-legal benchmarks into legal reliability scores.
AreaBenchmark or signalReported V4 Flash resultWhat the row can supportSource tag
Agentic codingSWE-bench Verified73.7% non-think to 79.0% Think Max [3]Strong software-engineering repair signal; mode-dependentProvider-reported DeepSeek table
Agentic terminal tasksTerminal-Bench 2.049.1% to 56.9% [3]Useful agentic execution signal; not a legal-reasoning measureProvider-reported DeepSeek table
Tool useToolathlon40.7% [3]Tool-use benchmark signal; does not test legal citation safetyProvider-reported DeepSeek table
MCP / agent contextMCP Atlas64% [3]Agentic workflow signalProvider-reported DeepSeek table
Agent evaluationClaw-Eval57.8% [3]Agent benchmark signalProvider-reported DeepSeek table
KnowledgeHumanity’s Last Exam8.1% non-think to 34.8% Think Max [3]Large mode effect; a non-think score is not interchangeable with a Think Max scoreProvider-reported DeepSeek table
Factual QASimpleQA23.1% [3]General factual-answering signal, not legal citation reliabilityProvider-reported DeepSeek table
Advanced knowledgeMMLU-Pro83.0% to 86.2% [3]General knowledge benchmark signal; mode label mattersProvider-reported DeepSeek table
Long-context retrieval / readingMRCR 1M37.5% to 78.7% [3]Very large mode-dependent gain; a 1M context window does not by itself prove stable legal recallProvider-reported DeepSeek table
Knowledge rankingBenchLM KnowledgeRank #55 of 55 [2]Adverse adjacent knowledge signalBenchLM model profile
Knowledge calibrationAA-Omniscience compilationApproximately 96% hallucination for Flash; approximately 94% for V4 Pro; Index -23 reported in adjacent summaries [4][5][6]Troubling calibration and abstention signal; exact per-model values should be checked against the live Artificial Analysis page during diligenceThird-party compilation / adjacent model overview; not a Flash legal benchmark
Legal benchmarkVals LegalBenchNo published V4 Flash row found; sibling V4 Pro appears at 80.32%, rank #71 of 129, snapshot July 23, 2026 [7]Missing Flash-specific legal-domain evidence; V4 Pro result is sibling evidence onlyBenchLM LegalBench mirror
Legal hallucination studyHalluHardV3.2-era DeepSeek family: 55% to 57% hallucinated legal claims without web search [8]Older-family adjacent warning; not a V4 Flash resultAcademic preprint

The strongest rows are the agentic-coding rows. SWE-bench Verified at 79.0% in Think Max, Terminal-Bench 2.0 up to 56.9%, and the supporting tool-use rows are exactly the kind of numbers that make V4 Flash credible as an engineering workhorse. They also explain why a legal-tech platform team would test it early rather than wait for a slower procurement cycle. [3]

But legal work does not only need an agent to complete steps. It needs the system to know when the record is insufficient, when authority is missing, and when a citation needs to be surfaced rather than invented. The AA-Omniscience signal is therefore not just another “knowledge benchmark.” A model that tends to answer when it does not know creates downstream work for the lawyer who has to check the answer, repair the memo, or explain why a false authority appeared in a draft. The approximately 96% Flash hallucination figure in the available compilation should be treated as a live diligence item, not a decorative red flag. [4][5]

Mode dependence is the other practical problem. On DeepSeek’s own reported tables, HLE moves from 8.1% in non-think mode to 34.8% in Think Max, and MRCR 1M moves from 37.5% to 78.7%. Those are not small tuning effects. If an evaluation memo reports “the V4 Flash score” without preserving the mode, the memo has already lost a material fact. [3]

The independent cross-check is useful, but it is not a Flash result

The best available model for reading DeepSeek’s provider-reported claims is CAISI/NIST’s May 1, 2026 evaluation of DeepSeek V4 Pro, a sibling model rather than V4 Flash. CAISI reported that some DeepSeek-reported results reproduced, including DeepSeek’s own GPQA-Diamond figure, but held-out independent benchmarks showed a capability lag of roughly eight months against the U.S. frontier. The report listed PortBench at 44% for V4 Pro versus 78% for the U.S. frontier, ARC-AGI-2 semi-private at 46% versus 79%, and IRT Elo 800 versus GPT-5.5’s 1260. [9]

NIST CAISI chart comparing DeepSeek V4 Pro overall AI capability on held-out benchmarks

That does not prove the same lag for V4 Flash. It does give procurement teams a disciplined way to read the dossier. Provider tables can be informative; they are not a substitute for held-out evaluation. Where provider claims reproduce, they become more credible. Where independent held-out benchmarks diverge, the procurement memo should say so in the capability column rather than hide the discrepancy in a footnote.

For Flash specifically, the inference should be labeled: sibling-model evidence suggests caution when translating DeepSeek’s reported benchmark strength into external capability, and Flash’s smaller active footprint may matter, but CAISI did not test Flash in that report. A legal risk score that imports the V4 Pro result directly into V4 Flash would be cleaner-looking than the evidence permits.

Vals LegalBench is the natural place to look for a legal-domain row. The public BenchLM mirror lists DeepSeek V4 Pro at 80.32%, rank #71 of 129, with a July 23, 2026 snapshot. It does not list V4 Flash. That distinction matters. V4 Pro’s result can inform a sibling-model inference; it cannot be quoted as Flash’s legal score. [7]

Even the available V4 Pro LegalBench row would not answer the full legal-risk question. Closed-book legal benchmarks can test legal knowledge and reasoning under benchmark conditions, but they do not establish that a model will provide accurate citations in a live filing workflow. Citation hallucination is a different failure mode from getting a multiple-choice or short-answer legal question right.

HalluHard adds a related warning, but it is also not a Flash result. The paper reports that a V3.2-era DeepSeek family produced hallucinated legal claims in the 55% to 57% range without web search. That is relevant to caution; it is not fair evidence that V4 Flash has the same legal hallucination rate. [8]

The court-record signal is also prospective rather than historical. Damien Charlotin’s database listed 1,811 AI hallucination cases as updated July 29, 2026, and the record materials for this dossier do not identify a court decision in that database naming DeepSeek V4 Flash. That absence should not be read as safety. It means there is no known sanction history tied to this specific model in that database snapshot. [10]

For legal teams, this is the uncomfortable middle position: there is enough evidence to justify a controlled technical pilot, and not enough evidence to let the model generate unsupervised legal authority. The missing Flash legal row belongs in the risk register as a positive fact: “not yet measured publicly,” not “assumed comparable to the coding rows.”

A useful V4 Flash pilot should not ask, in the abstract, whether the model is good. It should separate coding assistance, document operations, legal analysis, and citation production. Those are different tasks with different consequences.

  • Preserve the mode label for every run: non-think, Think, Think Max, or any internal wrapper mode used by the platform.
  • Record whether each result is provider-reported, internally reproduced, independently benchmarked, or inferred from a sibling model.
  • Do not count a correct-looking legal answer as citation-safe unless the cited authority is checked against a trusted source.
  • Route legal research through retrieval or grounding, and test whether the model can abstain when the source set does not contain the answer.
  • Separate cost testing from authority testing. A cheap successful extraction run does not validate legal reasoning.

The internal evaluation should end with a matrix, not a blended grade. “Approved for code generation inside the legal-tech platform” and “approved to draft cited legal propositions” should be separate decisions. A model can pass the first and fail the second.

Where the pilot touches filings, advice, or partner-facing research, use the site’s cite-verification workflow as the control plane, and compare any hallucination incident against the attorney sanctions digest. The benchmark score matters because it helps locate where verification labor will land.

Regulatory routing is separate from benchmark reliability

Some teams will never reach the benchmark question because policy or client constraints answer first. The Italy Garante ban, U.S. government-device rules, Chinese-model restrictions, and open-weight governance debates can affect whether V4 Flash is allowed into a workflow at all. Those questions should be routed through the site’s Chinese AI model regulation tracker and the open-weight restriction debate, not smuggled into a model-capability score.

Law-firm adoption signals are already conservative in some quarters. Bloomberg Law reported that Fox Rothschild blocked DeepSeek’s AI model for attorney use. That is not a benchmark result and does not prove V4 Flash is unreliable; it is a governance signal that peer firms may treat the model as a special approval item rather than an ordinary AI utility. [11]

For procurement purposes, the distinction is practical. A model may be capable enough to test and still be blocked by client terms, firm policy, government-device restrictions, or cross-border data rules. Conversely, a permitted deployment still needs a separate reliability record. Permission to use the tool is not proof that the tool can cite law safely.

Procurement verdict

Use caseV4 Flash fitReason
Agentic coding inside a legal-tech platformStrong candidate for controlled testingProvider-reported agentic-coding rows are materially strong and the API price is unusually low [1][3]
Document operations with supplied source textTestable with controlsLong context and agentic capability may help, but outputs still need task-specific QA [1][3]
Closed-book legal answeringNot approved on the public evidence recordNo published Flash legal benchmark row; adverse knowledge-calibration signals [4][5][7]
Citation-bearing legal researchNot approved without retrieval, grounding, and mandatory citation verificationLegal citation safety is not established by coding benchmarks or the sibling V4 Pro LegalBench row
Partner-facing or filing-adjacent draftingUse only as an assisted component with human review and verified sourcesKnown public evidence does not establish calibrated refusal or citation reliability for Flash

The clean conclusion is not that V4 Flash is bad. The record says something more useful: V4 Flash looks like an excellent low-cost agentic-coding model, a plausible workflow component, and a poor stand-alone legal authority. The danger is the blended scorecard that treats those conclusions as interchangeable.

Until a Flash-specific legal benchmark and citation-safety evaluation are published, the safe procurement label is: assisted workhorse only; legal authority supplied and checked elsewhere; all legal-risk estimates marked as inferred.

References

  1. DeepSeek V4 Preview Release, DeepSeek API Docs, April 24, 2026
  2. DeepSeek V4 Flash, BenchLM
  3. DeepSeek-V4-Flash, Hugging Face
  4. AI Hallucination Rates and Benchmarks, Suprmind
  5. Omniscience, Artificial Analysis
  6. DeepSeek V4 Pro Model Overview, DeepInfra
  7. Vals LegalBench, BenchLM, snapshot July 23, 2026
  8. HalluHard: A Legal Hallucination Benchmark, arXiv
  9. CAISI Evaluation of DeepSeek V4 Pro, NIST, May 1, 2026
  10. AI Hallucinations, Damien Charlotin, updated July 29, 2026
  11. Fox Rothschild Blocks DeepSeek’s AI Model for Attorney Use, Bloomberg Law

Chronological incident history

No sanction cases have named this tool in the tracked record set to date. This does not imply the tool is safe — see Risk Digest for ongoing monitoring.

← Compare peer tools

Report a correction or tip

Spotted an outdated figure, a misstated fact, or a ruling this tool profile should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.

Report a correction or tip for this record →
Blogarama - Blog Directory