Skip to content

Regulation

Who answers for AI drone strikes under the law of war?

By Editorial TeamUpdated Aug 2, 2026
Authority
International humanitarian law
Rule type
customary international law
Jurisdiction scope
International
Source text
Read primary rule text ↗

Document pre-deployment human decisions, target-class parameters, system evidence, and foreseeable electronic-warfare risks for AI-enabled drone strikes.

On June 1, 2025, Ukraine’s Operation Spider’s Web reportedly used roughly 100 drones against Russian airbases, including aircraft on the runway at Belaya air base. The law-of-war question that follows is not whether an AI-adjacent drone operation sits outside international humanitarian law. It does not. The harder question is narrower and more useful: when the strike process includes autonomous navigation, target recognition, terminal lock-on, or swarm-composition rules, whose decision can an investigator actually reconstruct afterward? [1]

Tu-22 bombers on the runway at Belaya air base during the June 1, 2025 drone strike

That matters for Ukraine drone strike law-of-war analysis because the operational record is moving faster than the courtroom record. CSIS reported, based on interviews, Ukrainian claims that AI-enabled autonomous navigation raised some strike success rates from roughly 10–20 percent to 70–80 percent, and that some systems could lock onto targets at about 2 kilometers. The same analysis noted that Ukraine lacked a formal definition of autonomy. Those are operationally significant claims, not adjudicated legal findings. They still describe the environment in which counsel, commanders, and investigators now have to work. [2]

Here, autonomy means a weapon system’s capacity, once activated, to perform some target-related function without a human making the final object-level choice at that moment. That includes modest forms such as autonomous navigation or terminal lock-on, and more consequential forms such as system selection among objects inside a human-defined target class. The definition is functional because legal and policy definitions diverge; the accountability problem appears before the definitional debate is settled.

IHL still applies; the weak point is attribution inside the targeting architecture

The baseline is not exotic. Distinction, proportionality, precautions in attack, weapons review duties, state responsibility, and individual criminal responsibility do not switch off because a drone uses machine perception or autonomous flight. For the fuller rule map, see the site’s baseline IHL obligations guide. The more difficult question is evidentiary: can prosecutors, defense counsel, military investigators, or reviewing commanders identify a foreseeable human decision connected to the unlawful harm?

The Lieber Institute’s July 2026 analysis puts the issue in useful terms: law-of-armed-conflict accountability depends on locating a foreseeable human decision, while autonomous-system architecture can diffuse that decision across commanders, operators, engineers, targeteers, data-labeling choices, model-performance limits, and pre-approved engagement parameters. The piece is expert legal analysis, not a court judgment or a state position, but its framing captures the practical problem better than the familiar claim that AI is simply “unaccountable.” [1]

A weapons lawyer does not need a new metaphysics of agency to see the problem. She needs to know who approved the target class, who selected the sensor package, who accepted the false-positive risk, who set the confidence threshold, who authorized reallocation inside the swarm, who received warnings about electronic-warfare vulnerability, and who had authority to abort. If those choices are undocumented, the legal system has not lost jurisdiction; it has lost the trail.

Three places the human decision can sit

Diagram comparing human target selection, human-defined target classes, and upstream swarm-composition rules

The Ukraine swarm-composition patterns discussed in the Lieber analysis are useful because they do not treat autonomy as a single switch. Responsibility looks different depending on where the decisive human judgment was placed before the system acted. [1]

When a human selects each target

The easiest case for later reconstruction is the one closest to traditional targeting. A human operator or commander selects a specific military objective, authorizes the strike, and the drone then uses AI-enabled navigation, obstacle avoidance, or terminal homing to reach the object. The system may still matter legally. If it cannot reliably distinguish the selected object from nearby protected persons or objects during the last meters of flight, that is not a trivial engineering detail. But the object-level decision remains visible: a human chose that target at that time.

In that pattern, the investigation starts with familiar questions. Was the selected object a military objective? What collateral-damage estimate was available? What alternatives were considered? What abort criteria existed if the object moved, the signal degraded, or civilians entered the area? What did the operator know about the drone’s terminal autonomy, and what was the basis for trusting it under the conditions present? The AI layer changes the evidence, but it does not hide the person whose target judgment must be assessed.

When a human defines the target class and the system chooses inside it

The harder pattern appears when a commander or targeteer does not select each individual object. Instead, the human defines a class: for example, a category of military equipment, a location boundary, a time window, and a set of visual or electronic signatures. The system then searches and selects among objects that appear to fit the class. In legal terms, the human decision has moved upstream. It is no longer “strike that object”; it is “strike objects that satisfy these parameters.”

That shift is not automatically unlawful. A target class can be narrow, militarily coherent, geographically bounded, and supported by reliable intelligence. It can also be overbroad, stale, poorly tested, or vulnerable to predictable misclassification. The legal review therefore has to examine the class definition itself: what objects were included, what objects were excluded, whether protected objects were likely to share the same signature, what confidence level was required, and whether the system was permitted to continue when sensor inputs became degraded.

This is where a surface-level approval record can mislead. A file may say that a human authorized a strike package against a lawful class of targets. That does not answer whether the class was defined with enough precision to preserve distinction, whether proportionality was assessed at a meaningful level, or whether precautions were tied to the system’s actual failure modes. The legal question is not merely whether a human clicked approve; it is whether the approval embodied the foreseeable judgment the law later needs to inspect.

When the decisive choice is embedded in swarm-composition rules

The most difficult pattern is not a single drone selecting a single object. It is a swarm in which upstream rules determine composition, routing, role assignment, target prioritization, decoy behavior, reallocation after losses, and conditions for continuing the attack. The human decision may be buried in the rule that says which drones carry which payloads, which sensor feeds govern target matching, how the swarm responds when communications are jammed, or which target categories receive priority when not all targets can be struck.

In that architecture, a commander may plausibly say that no human selected the specific object hit by Drone 27. That answer is incomplete. Someone approved the rule set that allowed Drone 27 to classify, prioritize, and engage under those conditions. Someone decided whether the swarm could substitute targets if the original target set disappeared. Someone decided whether loss of connectivity should trigger return, loiter, self-destruct, or continue-to-engage behavior. Someone accepted the risk that a decoy-detection rule would fail in a contested sensor environment.

This is the point at which command responsibility begins to strain. It can still attach to human knowledge, authority, and failure to prevent or punish. But if the military record treats the swarm-composition rule as a technical setting rather than as a targeting judgment, the later inquiry is forced to reconstruct legal responsibility from fragments: procurement records, test reports, model cards, mission-planning files, electronic-warfare assessments, operator training, and after-action data. The legal chain may exist; the documentary chain may not.

Why accountability becomes harder before the law changes

A compact diagnostic frame helps separate real accountability problems from loose rhetoric. Vincent Boulanin’s analysis identifies four reasons AI-driven autonomous weapons complicate legal accountability: unsettled application of existing rules, software unpredictability, diffusion of responsibility across many actors, and difficulty foreseeing system conduct at the moment the weapon is committed. [3]

Accountability pressure pointWhat it looks like in an AI-enabled drone strike
Unsettled applicationThe rule is familiar, but its application to target-class selection, autonomous reallocation, or terminal recognition has not been squarely resolved by courts.
Software unpredictabilityThe system may behave consistently with its training and design while still producing a result no operator specifically anticipated.
Diffusion across actorsA commander, operator, engineer, data curator, reviewer, and procurement authority may each control only part of the risk.
Foreseeability at commitmentThe legally important question becomes what conduct was reasonably foreseeable when the system was activated, not only what happened at impact.

The fourth point is the one that should make counsel uncomfortable. In traditional targeting, the commitment decision and the attack decision are often close enough in time and information that responsibility can be traced without too much conceptual strain. With autonomous swarms, the commitment decision may occur before the system encounters the decisive facts. The legal file must therefore preserve what the human decision-maker knew, and should have known, about how the system would behave when those facts appeared.

A tamper-evident log is necessary, and still not enough

Drone flight recorder with clean internal logs while electronic interference corrupts incoming sensor paths

The obvious answer is to demand logs. That answer is right as far as it goes. A tamper-evident audit record can preserve activation time, operator identity, mission parameters, software version, target-class settings, sensor inputs received, confidence scores, communications loss, abort signals, and the sequence leading to engagement. Without such records, later accountability depends on memory, inference, and whatever fragments survive the strike.

But the Lieber analysis flags the sharper problem: an internally coherent log may faithfully record corrupted inputs. In Ukraine’s electronic-warfare environment, adversarial interference can affect sensor feeds, navigation signals, communications, or classification cues. A drone may produce a clean record showing that it followed its rules, while the facts fed into those rules were distorted. A tamper-evident log can prove what the system believed; it does not by itself prove that the system’s belief was grounded in the external reality the law cares about. [1]

That distinction matters after a civilian object is hit. If the log says the system saw a valid target signature, the next question is how that signature was generated and verified. Did multiple sensors agree? Was one sensor known to be vulnerable to spoofing? Did the system continue after a confidence drop? Were there rules for cross-checking against no-strike data? Did the mission plan account for adversarial sensor corruption, or did it assume a benign input environment?

The audit record therefore has to connect three layers: the internal system state, the external sensor environment, and the human approval record. A log that records only machine events may help reconstruct causation but not responsibility. A legal review memo that records only human authorization may show compliance theater but not technical risk. The useful record links the two.

Command responsibility is doing the work, even if it does not fit cleanly

Human Rights Watch’s 2025 report presses a serious objection: command responsibility was built around human subordinates, while autonomous weapons may produce harm through system behavior that is not easily reducible to a subordinate’s criminal act. The report also warns that civil remedies in the United States may be blocked by Federal Tort Claims Act exceptions, including the discretionary-function and combatant-activities exceptions. [4]

That critique should not be brushed aside. It is strongest on remedies and evidentiary fit. If a victim cannot identify a human perpetrator, cannot obtain classified records, and cannot get past jurisdictional or immunity barriers, then a formal statement that “the law applies” may offer little practical accountability. The site’s prior record on drone-strike casualty remedies tracks some of those dead ends.

The critique is weaker if it is used to imply that commanders have no exposure until a future autonomous-weapons treaty arrives. Existing IHL still asks what commanders knew, what they should have known, what risks they accepted, and what measures they took to prevent or punish unlawful attacks. The pressure point is proving those facts when the decisive choices were made through target-class parameters, software settings, and swarm-composition rules rather than spoken strike orders.

Reported experience from Gaza’s Lavender targeting controversy is useful only as a cautionary counterpoint, not as a legal finding about Ukraine. Reports of high-volume targeting workflows, including alleged approvals in about 20 seconds and a reported error rate around 10 percent, show how human authorization can become thin when the system produces targets faster than humans can meaningfully review them. Those allegations remain theater-specific and contested; their relevance here is the institutional warning that “human in the loop” can become a label rather than a review function. [5]

Ukraine casualty figures should be handled with similar discipline. OHCHR civilian-casualty reporting provides documented minimum counts for the conflict and does not attribute casualties to AI autonomy. Those numbers may describe the scale of civilian harm in the war; they do not prove that autonomous drone functions caused particular incidents. [6]

What counsel should document before deployment

The practical duty is not to wait for treaty negotiations to settle every definition. It is to build the file that will let existing IHL operate after the strike. That file should be created before deployment, because the most important decisions may already be locked in once the swarm launches.

  • Record the human decision point: who approved the mission, who approved the target class, who approved autonomous reallocation, and who had abort authority.
  • Preserve parameter choices: target signatures, geographic limits, time windows, confidence thresholds, no-strike constraints, substitution rules, and communications-loss behavior.
  • Attach the system evidence: software version, model-performance testing, known failure modes, sensor dependencies, training limitations, and relevant updates since the last legal review.
  • Document foreseeable adversarial conditions: jamming, spoofing, decoys, degraded navigation, corrupted imagery, and how the system is supposed to behave when inputs conflict.
  • Make the audit record useful to lawyers, not only engineers: tie logs to mission approvals, target-class reasoning, precautionary measures, and post-strike review obligations.
  • Separate adoption from validation: the fact that a capability is being fielded does not prove it performs lawfully under the mission conditions in which it will be used.

Article 36-style review records should not be treated as a one-time certification if the relevant risk lies in mission configuration. A system that is lawful under one target-class definition, one sensor environment, and one communications plan may present different legal risks when those variables change. The same is true of a terminal-autonomy review. The site’s terminal-autonomy due-diligence workflow is the natural operational counterpart to this documentation file, and the Article 36 AI-drone review record covers the weapons-review side.

The central legal consequence is blunt. If the strike later causes unlawful harm, a commander cannot answer the accountability question by pointing to the autonomy of the system, and counsel cannot answer it by pointing to a generic human-approval step. The record must show where the foreseeable human judgment entered the architecture, what risks it accepted, and what evidence existed before the system was allowed to act.

References

  1. Whose Decision Was It? Drone Swarms and the Accountability Gap in Ukraine, Lieber Institute, July 24, 2026.
  2. Ukraine’s Future Vision and Current Capabilities for Waging AI-Enabled Autonomous Warfare, CSIS, March 6, 2025.
  3. Legal Accountability for AI-Driven Autonomous Weapons, Lieber Institute.
  4. A Hazard to Human Rights: Autonomous Weapons Systems and Digital Decision-Making, Human Rights Watch, April 28, 2025.
  5. Gaza Lavender reporting, Lieber, Mako & Al-Thani, May 4, 2026.
  6. Protection of Civilians in Armed Conflict — February 2026, OHCHR, February 2026.

Operationalizing workflow

No workflow has been explicitly linked to this obligation yet. See Workflows generally.

Illustrative cases

No illustrative case is currently tracked for this obligation. See Risk Digest for documented incidents generally.

← Back to Regulation

Report a correction or tip

Spotted an outdated figure, a misstated fact, or a ruling this regulation entry should reflect? Public comments are disabled for this content given the professional cost of a misreported case outcome, penalty amount, or rule text — use the structured correction channel instead.

Report a correction or tip for this record →