What Alibaba's Qwen Abstention Rate Means for Legal AI Tools
Alibaba's Qwen3.7 Max posts the lowest hallucination rate among comparable Chinese models by abstaining from roughly half of the questions it receives — which is why a low hallucination score does not, by itself, make a Qwen-based tool safe for legal research. This evaluation maps the model's legal-benchmark results to the workflows a Qwen-based tool can and cannot safely carry, so procurement and law-firm risk teams have a defensible answer on where it fits.
- Jurisdiction
- Global
- Court
- Various courts
- AI tool named
- Qwen3.7 Max
- Ruling date
- Jul 17, 2026
- Source document
- View primary court order ↗
- Last verified
- Aug 4, 2026
Lex Machina Review is an independent risk-tracking and reference resource. Nothing on this site is legal advice, and using it does not create an attorney-client relationship. Every record is reviewed against primary sources but may not reflect the most current status of a matter — always verify directly against the cited court order, rule text, or a licensed attorney before relying on it.
Companion explanation — secondary to the source document above
A low hallucination score is easy to misread in a procurement memo. Qwen3.7 Max, Alibaba’s frontier model, looks unusually careful on one broad hallucination benchmark because it often declines to answer. That matters for legal tools built on or benchmarked against Alibaba’s AI models: a refusal can be a useful safety behavior in a citation check, but it can also be a research failure when the lawyer needed the system to find missing authority.
The core numbers should be read together, not pulled apart into a single “hallucination rate” slide.
| Benchmark or source | Qwen3.7 Max result | Why it matters for legal-tool review |
|---|---|---|
| AA-Omniscience hallucination benchmark | 22.9% hallucination rate; 48% attempt rate; 30.1% accuracy [1] | The low hallucination figure comes with roughly half of questions not attempted. The figures are secondhand through an aggregator and should be traced to Artificial Analysis before use in a procurement file. |
| Vals AI Legal Research Bench | 25.48% ± 3.03; GPT-5.6 Sol is listed at 48.08%; DeepSeek V4 Pro at 23.08% on the same bench [2] | This is the more relevant comparison for legal research, and it does not support treating Qwen as a top legal-research performer. The live page should be re-verified before purchase review. |
| Harvey Legal Agent Benchmark | 1.67% [2] | A very low legal-agent score is hard to square with any claim that the model is ready to carry filing-dependent legal work without stronger tool-specific evidence. |

The procurement error is to treat the 22.9% hallucination rate as if it measures legal safety by itself. It does not. On the same reported AA-Omniscience snapshot, Qwen3.7 Max attempted 48% of the questions and reached 30.1% accuracy [1]. That is not a defect hidden in the footnotes; it is the mechanism. The model appears to be managing hallucination risk by refusing a large share of prompts rather than by answering nearly everything more accurately.
That behavior deserves credit. In a market where many models still produce fluent unsupported answers, visible abstention is a more honest failure mode. A system that says “I cannot answer from the materials provided” gives a lawyer something to triage. A system that fabricates a citation gives a lawyer something to apologize for.
But legal work is not scored only by avoiding false statements. In many workflows, the missing answer is the risk. If a system fails to surface a controlling case, a contrary regulation, or a procedural deadline, the fact that it did not hallucinate may not help the associate who relied on the search.
Abstention is useful only when the workflow can tolerate low recall
A Qwen-based legal tool can be defensible where the human already controls the relevant source set. If the user uploads a contract, a known statute, a case PDF, or a closed deal file, abstention can be made operationally useful. The model can be asked to extract defined clauses, check whether a quoted sentence appears in the document, compare a draft against provided source text, or identify whether a cited case name appears in a supplied appendix. If it refuses, the matter routes to a human reviewer. The cost is delay, not necessarily legal error.
The same profile is far less attractive for open-ended legal research. Comprehensive case retrieval, adverse-authority review, statutory interpretation across jurisdictions, and motion drafting from the open web or a full legal database are recall-heavy tasks. A model that declines half the time may look careful, but the lawyer still has to find the missing law somewhere else. In that setting, abstention does not complete the work; it transfers the work back to the legal team.

The distinction is not cosmetic. A citation-verification workflow can be designed around a narrow pass/fail result: found, not found, or escalate. A research workflow usually needs breadth. It needs the tool to locate relevant authorities, distinguish binding from persuasive material, flag negative treatment, and explain why omitted authority is not material. Low-confidence refusal may be preferable to fabrication, but it is not a substitute for recall.
For a knowledge-management lawyer, this means the acceptable deployment question is not “Does Qwen hallucinate less?” It is “What happens to the matter when Qwen does not answer?” If the answer is “a lawyer checks the same bounded source set,” the risk can be managed. If the answer is “the research may simply be incomplete,” the benchmark number is being used for the wrong decision.
The legal benchmarks do not support a broad research mandate
The Vals AI Legal Research Bench is the more relevant number for legal buyers than a general hallucination leaderboard. On that bench, Qwen3.7 Max is listed at 25.48% ± 3.03. The same snapshot lists GPT-5.6 Sol at 48.08% and DeepSeek V4 Pro at 23.08% [2]. That places Qwen near DeepSeek on this particular legal-research comparison and well behind the listed Western frontier model. It does not justify a conclusion that Qwen is unsafe in every legal workflow. It also does not justify a conclusion that Qwen’s low general hallucination rate makes it safe for legal research.
There is an important provenance caveat. The Vals and Harvey figures should be checked against the live Vals page before purchase review, because the figures were not confirmed in the captured page content and came from a search snippet [2]. That does not make the numbers unusable for analysis here; it means they should not be treated as procurement-grade evidence until the live source is captured and archived.
The Harvey Legal Agent Benchmark figure is even more sobering: Qwen3.7 Max is listed at 1.67% [2]. Without the full benchmark design and live verification, that number should not be overinterpreted. Still, a legal-agent benchmark score that low is not a basis for delegating agentic legal work. It is a reason to ask what the model was asked to do, what counted as success, and whether any Qwen-based product has independent legal-domain testing that is stronger than the base-model result.
Readers comparing frontier models on the same legal-research snapshot may also want the sibling evaluation, GPT-5.6 Leads DeepSeek V4 on Legal Research, Neither Is Safe. The useful comparison is not which model wins a headline; it is which workflow remains defensible after the model’s failure mode is understood.
Legal-specific systems have their own hallucination problem
The reason to be strict with Qwen is not that Chinese frontier models deserve a different standard. It is that legal AI systems, including purpose-built U.S. legal products, have already shown that legal-domain branding does not eliminate hallucination risk.
Stanford RegLab and Stanford HAI reported that, on a preregistered benchmark of more than 200 legal queries, Lexis+ AI hallucinated in more than 17% of responses and Westlaw AI-Assisted Research hallucinated in more than 34%. The same report said general-purpose chatbots hallucinated on legal queries at rates from 58% to 82% [3]. Those figures are not a direct comparison to Qwen3.7 Max’s AA-Omniscience score; they are a warning against moving a model across benchmarks and pretending the risk measure stayed the same.
The warning is reinforced by broader hallucination testing that identifies legal information as a particularly difficult knowledge domain: an 18.7% average hallucination rate across models, compared with 6.4% for top models in the same comparison [1]. Again, the proper lesson is narrow. Legal work is a high-friction domain for language models, even before a buyer asks the system to perform retrieval, cite authority, or support a filing.
That is why the sanction environment matters. This site’s current case-tracking context records 1,769 documented AI-hallucination rulings globally as of July 17, 2026, with Q1 2026 sanctions exceeding $145,000; for the Western-frontier baseline and case-count discussion, see How Reliable Is ChatGPT for Legal Work in 2026?. In that environment, a tool’s refusal behavior is not merely a UX feature. It affects who signs the filing, who reviews the output, and who explains the workflow if a hallucinated or omitted authority reaches the court.
Why Qwen still matters to legal AI buyers
There is no need to pretend that Qwen is irrelevant to legal tools merely because the public record does not identify a named U.S.-market legal product built on Qwen. Alibaba’s model family has broad open-weight reach. Wikipedia’s Qwen entry describes Qwen as having more than 1 billion downloads and accounting for roughly 40% of newly created Hugging Face models [4]. Those figures are not proof of U.S. legal-product adoption. They do explain why risk teams may encounter Qwen-based derivatives, internal prototypes, vendor demos, or fine-tuned legal models even if the commercial product name does not advertise the base model.
That is the correct frame for Qwen-based legal models such as Qzhou-Law and Alibaba’s own Tongyi Farui: they make the model-family question relevant, but they do not prove that a U.S. law firm is already buying a Qwen-backed research platform. A procurement file should not fill that gap with implication. It should ask the vendor directly which base model is used, whether the legal layer changes refusal behavior, what retrieval system supplies authority, and whether the legal benchmark has been independently reproduced.
The same discipline applies to agentic claims. Vendor-stated autonomous-operation or coding-benchmark results may be interesting engineering signals, but they are not legal-research evidence without independent reproduction and a legal task definition. A model that performs well in a software or terminal benchmark has not thereby shown that it can find binding authority, preserve citation integrity, or avoid filing-risk hallucinations.
A defensible placement for Qwen-based legal tools
A risk committee does not need a philosophical position on Alibaba, Chinese models, or open weights. It needs a deployment boundary that can survive scrutiny. Based on the reported profile, the safer boundary is narrow: use a Qwen-based tool only where the workflow is verification-heavy, source-bounded, and designed to escalate refusals.
- More defensible: checking whether a known citation appears in supplied materials; comparing a draft sentence against uploaded authority; extracting clauses from a closed contract set; summarizing documents where the human reviewer already controls the record.
- Less defensible: comprehensive legal research; adverse-authority searches; open-ended jurisdictional analysis; drafting briefs or memos that depend on finding law outside the provided record.
- Not procurement-grade without stronger evidence: filing-dependent legal-agent work, unsupervised research routing, or workflows where a refusal silently narrows the legal universe reviewed.
The audit question should be concrete. If the model refuses, is the refusal logged? Does a human receive the task? Does the matter stop, reroute, or proceed with an incomplete answer? Does the user see that the answer is based only on provided materials? Those controls matter more than a generic statement that the model has a low hallucination rate.
Qwen3.7 Max’s abstention strategy is a meaningful reliability signal, but it is not a legal-safety certificate. It supports cautious use in bounded verification and draft-from-source workflows under human supervision. It does not support deployment for high-recall legal research, open-ended legal analysis, or filing-dependent work unless a Qwen-based legal product can show stronger, independently traceable, legal-specific evidence.
References
- AI Hallucination Rates and Benchmarks, Suprmind
- Alibaba Qwen3.7 Max, Vals AI
- AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries, Stanford HAI
- Qwen, Wikipedia
Related records
Tool profile
Browse tool evaluations →Governing regulation
Browse the obligations tracker →Preventive workflow
Browse verification workflows →
Report a correction or tip
Spotted an outdated figure, a misstated fact, or a ruling this case record should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.
Report a correction or tip for this record →