DeepSeek V4 Flash agent benchmark scorecard for legal teams
DeepSeek V4 Flash posts class-leading agentic-coding scores at outlier-low prices, yet has no published legal-domain benchmark and shows near-record overconfidence on knowledge tasks. This scorecard separates provider-reported from independent results and shows legal teams exactly where V4 Flash needs a grounding and cite-verification layer before use.
- Tool
- DeepSeek V4 Flash
- Benchmark source
- DeepSeek provider tables; BenchLM; Artificial Analysis; CAISI/NIST
- Hallucination rate
- approx. 96% (AA-Omniscience)
- Test methodology
- Mixed benchmark ledger separating provider-reported rows from independent/adjacent evidence, with thinking-mode labels and sibling-model inferences flagged; no Flash-specific legal benchmark.
- Test date
- Jul 31, 2026
DeepSeek V4 Flash is the kind of model that makes an engineering team pause: an open-weight MoE model, a 1M-token context window, MIT licensing, and API pricing listed at $0.14 input and $0.28 output per 1M tokens. It also posts very strong provider-reported agentic-coding results. The legal problem is that the evidence gets thin exactly where a legal team needs it to get specific: legal authority, citation reliability, and calibrated refusal. As of this record date, the public record contains no Flash-specific legal-domain benchmark row, while available knowledge-calibration evidence points to a model that may answer far too often when it should abstain. [1][2][3][4][5]

The immediate procurement answer is therefore narrow. V4 Flash is worth testing as a low-cost coding, agent, and workflow component. It should not be treated as a legal authority unless the legal authority is supplied by a retrieval or grounding layer and the citations are independently verified. Any legal-risk estimate for Flash is inferred from sibling-model and adjacent evidence, not measured directly on a published Flash legal benchmark.
Model record
| Field | Record for DeepSeek-V4-Flash | Evidence label |
|---|---|---|
| Release | April 24, 2026 [1] | Provider release note |
| Architecture footprint | 284B total parameters / 13B active MoE [2] | BenchLM model profile |
| Context length | 1M tokens [1][2] | Provider release note and model profile |
| License | MIT [2] | BenchLM model profile |
| API price | $0.14 input / $0.28 output per 1M tokens [1] | Provider release note |
| Procurement status for legal work | Testable as an assisted workhorse; not acceptable as a stand-alone legal authority without grounding and cite verification | Evaluation judgment from the benchmark record |
That price-performance profile is not a side detail. At legal-department scale, low inference cost changes which experiments are affordable: bulk document routing, coding-heavy integrations, contract-data extraction pipelines, and agentic testing harnesses all become easier to justify. The mistake would be to let that operational attractiveness turn every benchmark row into a legal-reliability row.
Benchmark ledger: what the scores actually measure
A single benchmark table is not enough for this model. V4 Flash’s profile changes by task type and, in several rows, by reasoning mode. A legal buyer needs the label before the score: provider-reported or independent, coding or knowledge, closed-book or grounded, non-think or Think Max.

| Area | Benchmark or signal | Reported V4 Flash result | What the row can support | Source tag |
|---|---|---|---|---|
| Agentic coding | SWE-bench Verified | 73.7% non-think to 79.0% Think Max [3] | Strong software-engineering repair signal; mode-dependent | Provider-reported DeepSeek table |
| Agentic terminal tasks | Terminal-Bench 2.0 | 49.1% to 56.9% [3] | Useful agentic execution signal; not a legal-reasoning measure | Provider-reported DeepSeek table |
| Tool use | Toolathlon | 40.7% [3] | Tool-use benchmark signal; does not test legal citation safety | Provider-reported DeepSeek table |
| MCP / agent context | MCP Atlas | 64% [3] | Agentic workflow signal | Provider-reported DeepSeek table |
| Agent evaluation | Claw-Eval | 57.8% [3] | Agent benchmark signal | Provider-reported DeepSeek table |
| Knowledge | Humanity’s Last Exam | 8.1% non-think to 34.8% Think Max [3] | Large mode effect; a non-think score is not interchangeable with a Think Max score | Provider-reported DeepSeek table |
| Factual QA | SimpleQA | 23.1% [3] | General factual-answering signal, not legal citation reliability | Provider-reported DeepSeek table |
| Advanced knowledge | MMLU-Pro | 83.0% to 86.2% [3] | General knowledge benchmark signal; mode label matters | Provider-reported DeepSeek table |
| Long-context retrieval / reading | MRCR 1M | 37.5% to 78.7% [3] | Very large mode-dependent gain; a 1M context window does not by itself prove stable legal recall | Provider-reported DeepSeek table |
| Knowledge ranking | BenchLM Knowledge | Rank #55 of 55 [2] | Adverse adjacent knowledge signal | BenchLM model profile |
| Knowledge calibration | AA-Omniscience compilation | Approximately 96% hallucination for Flash; approximately 94% for V4 Pro; Index -23 reported in adjacent summaries [4][5][6] | Troubling calibration and abstention signal; exact per-model values should be checked against the live Artificial Analysis page during diligence | Third-party compilation / adjacent model overview; not a Flash legal benchmark |
| Legal benchmark | Vals LegalBench | No published V4 Flash row found; sibling V4 Pro appears at 80.32%, rank #71 of 129, snapshot July 23, 2026 [7] | Missing Flash-specific legal-domain evidence; V4 Pro result is sibling evidence only | BenchLM LegalBench mirror |
| Legal hallucination study | HalluHard | V3.2-era DeepSeek family: 55% to 57% hallucinated legal claims without web search [8] | Older-family adjacent warning; not a V4 Flash result | Academic preprint |
The strongest rows are the agentic-coding rows. SWE-bench Verified at 79.0% in Think Max, Terminal-Bench 2.0 up to 56.9%, and the supporting tool-use rows are exactly the kind of numbers that make V4 Flash credible as an engineering workhorse. They also explain why a legal-tech platform team would test it early rather than wait for a slower procurement cycle. [3]
But legal work does not only need an agent to complete steps. It needs the system to know when the record is insufficient, when authority is missing, and when a citation needs to be surfaced rather than invented. The AA-Omniscience signal is therefore not just another “knowledge benchmark.” A model that tends to answer when it does not know creates downstream work for the lawyer who has to check the answer, repair the memo, or explain why a false authority appeared in a draft. The approximately 96% Flash hallucination figure in the available compilation should be treated as a live diligence item, not a decorative red flag. [4][5]
Mode dependence is the other practical problem. On DeepSeek’s own reported tables, HLE moves from 8.1% in non-think mode to 34.8% in Think Max, and MRCR 1M moves from 37.5% to 78.7%. Those are not small tuning effects. If an evaluation memo reports “the V4 Flash score” without preserving the mode, the memo has already lost a material fact. [3]
The independent cross-check is useful, but it is not a Flash result
The best available model for reading DeepSeek’s provider-reported claims is CAISI/NIST’s May 1, 2026 evaluation of DeepSeek V4 Pro, a sibling model rather than V4 Flash. CAISI reported that some DeepSeek-reported results reproduced, including DeepSeek’s own GPQA-Diamond figure, but held-out independent benchmarks showed a capability lag of roughly eight months against the U.S. frontier. The report listed PortBench at 44% for V4 Pro versus 78% for the U.S. frontier, ARC-AGI-2 semi-private at 46% versus 79%, and IRT Elo 800 versus GPT-5.5’s 1260. [9]

That does not prove the same lag for V4 Flash. It does give procurement teams a disciplined way to read the dossier. Provider tables can be informative; they are not a substitute for held-out evaluation. Where provider claims reproduce, they become more credible. Where independent held-out benchmarks diverge, the procurement memo should say so in the capability column rather than hide the discrepancy in a footnote.
For Flash specifically, the inference should be labeled: sibling-model evidence suggests caution when translating DeepSeek’s reported benchmark strength into external capability, and Flash’s smaller active footprint may matter, but CAISI did not test Flash in that report. A legal risk score that imports the V4 Pro result directly into V4 Flash would be cleaner-looking than the evidence permits.
The legal-benchmark gap is a finding, not an empty cell
Vals LegalBench is the natural place to look for a legal-domain row. The public BenchLM mirror lists DeepSeek V4 Pro at 80.32%, rank #71 of 129, with a July 23, 2026 snapshot. It does not list V4 Flash. That distinction matters. V4 Pro’s result can inform a sibling-model inference; it cannot be quoted as Flash’s legal score. [7]
Even the available V4 Pro LegalBench row would not answer the full legal-risk question. Closed-book legal benchmarks can test legal knowledge and reasoning under benchmark conditions, but they do not establish that a model will provide accurate citations in a live filing workflow. Citation hallucination is a different failure mode from getting a multiple-choice or short-answer legal question right.
HalluHard adds a related warning, but it is also not a Flash result. The paper reports that a V3.2-era DeepSeek family produced hallucinated legal claims in the 55% to 57% range without web search. That is relevant to caution; it is not fair evidence that V4 Flash has the same legal hallucination rate. [8]
The court-record signal is also prospective rather than historical. Damien Charlotin’s database listed 1,811 AI hallucination cases as updated July 29, 2026, and the record materials for this dossier do not identify a court decision in that database naming DeepSeek V4 Flash. That absence should not be read as safety. It means there is no known sanction history tied to this specific model in that database snapshot. [10]
For legal teams, this is the uncomfortable middle position: there is enough evidence to justify a controlled technical pilot, and not enough evidence to let the model generate unsupervised legal authority. The missing Flash legal row belongs in the risk register as a positive fact: “not yet measured publicly,” not “assumed comparable to the coding rows.”
What a legal pilot should preserve
A useful V4 Flash pilot should not ask, in the abstract, whether the model is good. It should separate coding assistance, document operations, legal analysis, and citation production. Those are different tasks with different consequences.
- Preserve the mode label for every run: non-think, Think, Think Max, or any internal wrapper mode used by the platform.
- Record whether each result is provider-reported, internally reproduced, independently benchmarked, or inferred from a sibling model.
- Do not count a correct-looking legal answer as citation-safe unless the cited authority is checked against a trusted source.
- Route legal research through retrieval or grounding, and test whether the model can abstain when the source set does not contain the answer.
- Separate cost testing from authority testing. A cheap successful extraction run does not validate legal reasoning.
The internal evaluation should end with a matrix, not a blended grade. “Approved for code generation inside the legal-tech platform” and “approved to draft cited legal propositions” should be separate decisions. A model can pass the first and fail the second.
Where the pilot touches filings, advice, or partner-facing research, use the site’s cite-verification workflow as the control plane, and compare any hallucination incident against the attorney sanctions digest. The benchmark score matters because it helps locate where verification labor will land.
Regulatory routing is separate from benchmark reliability
Some teams will never reach the benchmark question because policy or client constraints answer first. The Italy Garante ban, U.S. government-device rules, Chinese-model restrictions, and open-weight governance debates can affect whether V4 Flash is allowed into a workflow at all. Those questions should be routed through the site’s Chinese AI model regulation tracker and the open-weight restriction debate, not smuggled into a model-capability score.
Law-firm adoption signals are already conservative in some quarters. Bloomberg Law reported that Fox Rothschild blocked DeepSeek’s AI model for attorney use. That is not a benchmark result and does not prove V4 Flash is unreliable; it is a governance signal that peer firms may treat the model as a special approval item rather than an ordinary AI utility. [11]
For procurement purposes, the distinction is practical. A model may be capable enough to test and still be blocked by client terms, firm policy, government-device restrictions, or cross-border data rules. Conversely, a permitted deployment still needs a separate reliability record. Permission to use the tool is not proof that the tool can cite law safely.
Procurement verdict
| Use case | V4 Flash fit | Reason |
|---|---|---|
| Agentic coding inside a legal-tech platform | Strong candidate for controlled testing | Provider-reported agentic-coding rows are materially strong and the API price is unusually low [1][3] |
| Document operations with supplied source text | Testable with controls | Long context and agentic capability may help, but outputs still need task-specific QA [1][3] |
| Closed-book legal answering | Not approved on the public evidence record | No published Flash legal benchmark row; adverse knowledge-calibration signals [4][5][7] |
| Citation-bearing legal research | Not approved without retrieval, grounding, and mandatory citation verification | Legal citation safety is not established by coding benchmarks or the sibling V4 Pro LegalBench row |
| Partner-facing or filing-adjacent drafting | Use only as an assisted component with human review and verified sources | Known public evidence does not establish calibrated refusal or citation reliability for Flash |
The clean conclusion is not that V4 Flash is bad. The record says something more useful: V4 Flash looks like an excellent low-cost agentic-coding model, a plausible workflow component, and a poor stand-alone legal authority. The danger is the blended scorecard that treats those conclusions as interchangeable.
Until a Flash-specific legal benchmark and citation-safety evaluation are published, the safe procurement label is: assisted workhorse only; legal authority supplied and checked elsewhere; all legal-risk estimates marked as inferred.
References
- DeepSeek V4 Preview Release, DeepSeek API Docs, April 24, 2026
- DeepSeek V4 Flash, BenchLM
- DeepSeek-V4-Flash, Hugging Face
- AI Hallucination Rates and Benchmarks, Suprmind
- Omniscience, Artificial Analysis
- DeepSeek V4 Pro Model Overview, DeepInfra
- Vals LegalBench, BenchLM, snapshot July 23, 2026
- HalluHard: A Legal Hallucination Benchmark, arXiv
- CAISI Evaluation of DeepSeek V4 Pro, NIST, May 1, 2026
- AI Hallucinations, Damien Charlotin, updated July 29, 2026
- Fox Rothschild Blocks DeepSeek’s AI Model for Attorney Use, Bloomberg Law
Chronological incident history
No sanction cases have named this tool in the tracked record set to date. This does not imply the tool is safe — see Risk Digest for ongoing monitoring.
← Compare peer toolsReport a correction or tip
Spotted an outdated figure, a misstated fact, or a ruling this tool profile should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.
Report a correction or tip for this record →