How to Assess Terminal Autonomy Liability in Drone Procurement
A structured due-diligence workflow for legal advisors evaluating terminal-phase autonomy architectures in drone systems under IHL accountability doctrines, using the three-pattern framework from recent operational analysis to identify whether a system's design precludes meaningful human control at the engagement point.
- Applicable role
- attorney
- Workflow stage
- pre-filing
- Primary source
- DoD Directive 3000.09 (2023 update)
The procurement question is not whether the proposal uses the word “autonomous.” The question is narrower and more uncomfortable: when an AI-guided drone enters terminal guidance, where does meaningful human control structurally end? If the system can continue to identify, select, and engage a target after communications with the operator are severed, the legal file has to show more than a clean approval slide. It has to show which human decision, at which time, can carry distinction, proportionality, and precaution.
There is no published U.S. court opinion squarely resolving terminal-autonomy liability for AI-guided drones in the setting that matters most here: a weapon system that executes the final engagement autonomously after the human can no longer intervene. That leaves counsel with a risk-assessment task rather than a precedent checklist. The most useful starting point is the Lieber Institute’s July 24, 2026 analysis of drone swarms and accountability in Ukraine, because it translates the debate into three inspectable terminal-autonomy patterns: individual-threshold, distributed-success, and leader-follower architectures.[1]

Terminal Autonomy Is a Commitment-Point Problem
Terminal autonomy should be kept narrow. It is not every onboard navigation function, every automated sensor process, or every loss-of-link routine. The relevant moment is the commitment point: the phase in which the system moves from search, track, or route adjustment into an engagement that a human can no longer practically stop.
That distinction matters because ordinary procurement language often sits upstream. A file may say that a commander approved the mission, that an operator initiated launch, that the system was tested against approved target profiles, or that the human retains supervisory control. Those statements may be true and still miss the liability signal if the architecture removes the human before the target is finally selected or before the attack becomes irreversible.
The legal review therefore starts by asking engineers to draw the control chain. Who receives sensor information? Who or what classifies the object? Who sets the threshold for engagement? What happens when the communications link is degraded or severed? Is there an abort channel during terminal descent? If there is no abort channel, is the remaining action merely ballistic execution of a prior human decision, or is the system still making target-relevant judgments?
The Three Architecture Patterns That Change the Accountability Question
The Lieber Institute framework is useful because it does not ask counsel to decide in the abstract whether “autonomy” is lawful. It asks where the terminal decision is composed. The answer changes the due-diligence file.

| Terminal-autonomy pattern | Where the terminal decision forms | Primary due-diligence question |
|---|---|---|
| Individual-threshold | Inside a single drone after its onboard system reaches a configured confidence or engagement threshold | Who set the threshold, what target class did it encode, and could a human still stop the engagement after the threshold was met? |
| Distributed-success | Across multiple drones or nodes that share sensing, target confirmation, or engagement roles | Can the record reconstruct which node supplied the target-relevant judgment and which human approved the combined effect? |
| Leader-follower | In a lead drone or lead element that directs follower drones once the human link is absent or attenuated | Is the legally relevant decision in the human’s mission order, the leader system’s terminal choice, or the followers’ execution logic? |
Individual-Threshold Systems
In an individual-threshold architecture, a single drone carries the decisive terminal logic. The operator may select an area, launch the platform, approve a mission profile, or authorize a class of targets. But the engagement occurs only when the drone’s onboard system decides that sensor inputs satisfy a threshold configured before the terminal phase.
That pattern is not automatically unlawful. It is, however, unforgiving from an accountability perspective. If the human decision is “engage only military vehicles of a specified class,” the file has to show how the system recognizes that class, how uncertainty is handled, what happens at the edge cases, and whether the operator can refuse engagement after the onboard threshold is met. If the system crosses the threshold after communications are severed, the earlier human approval must be precise enough to bear the legal weight of the later strike.
The hard case is not a malfunction in the casual sense. It is normal operation that leaves no person at the terminal engagement point. In that setting, post-strike review may identify what the software did, but it may not identify a human who made the distinction judgment when the relevant information became available.
Distributed-Success Systems
Distributed-success architectures create a different evidentiary problem. No single drone may appear to make the whole decision. One node detects movement, another confirms a pattern, another relays location, and another executes. The engagement looks like the product of a system rather than a person or even a single machine.
For procurement review, that means logs are not a technical afterthought. They are the only way to test whether the legal decision can be reconstructed. Counsel should want to know whether the system records which node contributed which classification, whether confidence scores or equivalent decision variables are retained, whether the operator display showed fused conclusions or raw disagreement among nodes, and whether the architecture permits a human to veto the combined result before engagement.
A distributed system can diffuse responsibility without anyone intending to hide it. That is exactly why the procurement file should not accept “human supervised” as a complete answer. Supervision of a swarm-level output is not the same as a human decision on the target that was actually struck.
Leader-Follower Systems
Leader-follower systems look cleaner on paper because there is a visible hierarchy. A lead drone or lead element identifies, designates, or routes; follower drones act in response. But the legal question remains: who made the terminal engagement decision, and at what point could a human still say no?
If the lead platform merely transmits a human-selected target to followers, the accountability chain may remain tied to the human who authorized the attack. If the lead platform continues to select or reprioritize targets after the operator link is gone, the decision has moved. The followers may be executing faithfully, while the legally relevant judgment sits in the lead system’s autonomous terminal logic.
This is where diagrams beat adjectives. “Semi-autonomous,” “supervised,” and “operator controlled” can all be accurate at different phases. They do not answer whether the lead system can create a new engagement condition after the human has lost practical control.
Policy Language Cannot Substitute for the Control Chain
DoD Directive 3000.09 remains a central procurement reference for autonomy in weapon systems. Its 2023 update is also a reminder that small wording choices can carry oversight consequences. Legal commentary has flagged a contested shift from “human” oversight language toward “operator” oversight language; some scholars read that change as potentially broad enough to permit non-human or AI-enabled oversight roles, while others view it as administrative tightening rather than a substantive relaxation of review obligations.[2]
The safe procurement response is not to turn that dispute into a slogan. It is to define “operator” functionally in the file. Is the operator a person who can perceive the relevant target information, understand the system’s recommendation, and stop the engagement during the terminal phase? Or is the operator a role named in the documentation while the architecture places the final engagement beyond human refusal?
That distinction should appear in requirements, test plans, and acceptance criteria. A vendor representation that the system keeps an operator “on the loop” has limited value unless the file shows what information reaches that person, how much time remains for refusal, what the abort path is, and whether the system can continue to engage when that path is unavailable.
Command Responsibility Does Not Disappear, but It May Not Solve the Whole Problem
The ICC arrest warrants issued in 2024 for Sergei Kobylash, Sergei Shoigu, and Valery Gerasimov are a useful warning for anyone tempted to treat autonomy as a command shield. The warrants, discussed in current legal analysis of autonomous drone accountability, concern alleged command responsibility connected to Shahed-136 strikes against Ukrainian civilian infrastructure.[1][3]
They should not be made to do more than they can do. The warrants remain unexecuted as of July 31, 2026, and they do not supply a tested rule for attributing responsibility to programmers, designers, or procurement officials. They do show that command responsibility remains part of the accountability landscape. They do not answer the procurement lawyer’s architecture question: did the system design leave a human commander or operator with a meaningful decision at the engagement point?
Temple iLIT’s analysis frames the related three-actor problem: operator, commander, and designer or programmer. Terminal autonomy complicates attribution because intent, knowledge, and foreseeability may be split across those roles. The operator may not know what the model will classify in the terminal phase. The commander may approve an operation without knowing how much discretion has shifted to the system. The designer may foresee classes of failure without knowing the operational context in which the system will be used.[3]
Procurement counsel should treat that split as a documentable risk, not an after-action surprise. If the contract record cannot identify which actor owns the relevant judgment, the answer is not improved by waiting until the strike has occurred.

What the Procurement File Should Prove
A defensible review does not need to solve every future battlefield contingency. It does need to show that the buyer understood where the terminal decision sits and what that placement does to legal accountability. The file should answer a few concrete questions before approval.
- Commitment point: Identify the last point at which a person can stop the engagement, not merely the last point at which a person approved the mission.
- Target-relevant information: Record what the human sees before that point, including whether the display shows raw sensor data, a system classification, a confidence indicator, or only a recommended action.
- Autonomous discretion: State whether the system may select, reprioritize, or confirm a target after communications are severed.
- Abort architecture: Describe the technical means for refusal or abort during terminal guidance, and the system behavior when that means fails.
- Evidence retention: Require logs that can reconstruct sensor inputs, classification steps, threshold crossings, node-to-node messages, operator prompts, and abort commands.
- Human legal attachment: Name the person or role whose decision is expected to carry distinction, proportionality, and precaution at the relevant time.
The last item is the one most likely to be softened in procurement language. It should not be. A system can have excellent performance metrics and still create legal exposure if the legal judgment is assigned to a person who never receives the terminal facts or has no practical ability to refuse the engagement.
Distinction
For distinction, the review should track the object or person classification chain. Who determined that the target class was military? Who decided that the sensor features used by the system were adequate proxies for that class? If the drone makes the final classification after link loss, what pre-deployment decision is supposed to stand in for the missing human judgment?
The answer may be stronger for tightly bounded target sets and weaker for open-ended environments. That is not a moral intuition; it is an evidentiary point. The broader the environment and the more ambiguous the target signature, the harder it becomes to show that a prior human decision actually resolved the distinction question presented at terminal engagement.
Proportionality
For proportionality, the file should not stop at target value. It should ask whether the person making the proportionality judgment had information about expected civilian harm at the time the attack became irreversible. If the system can update or choose a target after that judgment, counsel should ask whether the original proportionality assessment still corresponds to the actual engagement.
This is particularly difficult in distributed-success systems. If one node contributes location, another contributes classification, and another executes, the proportionality record may show mission-level approval without showing that anyone assessed the combined, final strike condition.
Precaution
For precaution, the question is whether feasible precautions remain feasible by design. A paper procedure requiring human review is thin protection if the terminal phase gives the operator no usable information, no time, or no technical channel to act on it. The file should show what precautionary measures are embedded before launch and what measures remain available during terminal guidance.
A procurement record that treats link loss only as a reliability issue is incomplete. In terminal autonomy, link loss can also be the moment when the legal review function becomes historical.
Foreseeability Belongs in the Design Review
The recurring difficulty in autonomous-weapon accountability is not merely that something can go wrong. It is that responsibility can be spread across unsettled law, software behavior, multiple human and institutional actors, and system conduct that may be difficult to foresee at the commitment point. Those are not reasons to abandon review. They are reasons to move the review earlier.
Foreseeability in procurement should be framed at the level of architecture. It is foreseeable that an individual-threshold system may engage after the operator loses contact if the requirements permit exactly that. It is foreseeable that a distributed-success system may make attribution difficult if no node-level decision record is retained. It is foreseeable that a leader-follower system may obscure the legal decision if the lead system can generate target choices for followers outside human control.
That does not mean every later error was specifically predicted. It means the accountability gap was not accidental if the approved design placed the terminal engagement beyond human refusal and failed to preserve evidence of how the decision formed.
Regulatory Direction Is Moving Toward Architecture-Based Limits
The broader regulatory direction reinforces the same due-diligence instinct. The UN Secretary-General’s 2026 call for a legally binding instrument reflects a dual-track approach: prohibit systems incapable of complying with international humanitarian law and regulate systems with partial autonomy. That distinction puts pressure on procurement records to show not only that a system performs, but that its design preserves legally meaningful control where IHL requires judgment.
For counsel, the practical implication is modest but important. Do not wait for treaty language, national implementing rules, or litigation to force the architecture question into the file. If the system’s terminal design already makes human judgment unavailable at engagement, the risk is present even before a court or regulator supplies a final label.
The Due-Diligence Conclusion
A legal review of terminal autonomy should end with a design finding, not a branding finding. The file should state which terminal-autonomy pattern the system uses, where the decision break point occurs, what human decision remains at that point, and what evidence will prove it after the fact.
If the architecture preserves a human ability to assess the target, weigh expected harm, take feasible precautions, and refuse the engagement before it becomes irreversible, the procurement record can say so and identify the proof. If the architecture does not preserve that ability, the record should not bury the issue under supervision language. The legal risk is foreseeable when the system design itself precludes meaningful human control at engagement.
That is the central liability signal in drone procurement. The problem is not discovered only after a disputed strike. It is visible when the control chain is drawn.
References
- Whose Decision Was It? Drone Swarms and the Accountability Gap in Ukraine, Lieber Institute at West Point, July 24, 2026
- Autonomy in Weapon Systems, U.S. Department of Defense, 2023
- Lethal Autonomous Weapon Systems (LAWS): Accountability, Collateral Damage, and the Inadequacies of International Law, Temple iLIT
Grounded in
This procedure is grounded in DoD Directive 3000.09 (2023 update), independent of any single documented case. See the Regulation tracker for the governing text.
Cases this step would have prevented
No cases have been explicitly linked to this checklist yet. See Risk Digest for documented incidents generally.
← Back to WorkflowsReport a correction or tip
Spotted an outdated figure, a misstated fact, or a ruling this workflow checklist should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.
Report a correction or tip for this record →