When AI Soldiers Follow Orders, Who Refuses the Unlawful Ones?

Pentagon’s pursuit of low-refusal AI for national security leaves human accountability for unlawful orders without a clear owner

Annemarije de Boer Avatar
Annemarije de Boer Avatar

By

Image: Public Domain Pictures – Jean Beaufort

Key Takeaways

Key Takeaways

  • Pentagon draft described OpenAI models for national security with “minimal refusal rates.”
  • Soldiers must refuse unlawful orders; AI refusal is a technical control, not legal duty.
  • Anthropic’s Pentagon exclusion highlights risks of removing AI safeguards for classified contracts.

A Pentagon contract draft described OpenAI Mission Models” as systems designed for national-security use cases with “minimal refusal rates.” That phrase, obtained by The Intercept through a Freedom of Information Act lawsuit, sounds like a performance specification. It raises a harder question: when an AI system is optimized to comply, what happens to the safeguard that exists precisely for the moment a command is wrong?

A Disputed Phrase and an Unresolved Record

The “minimal refusal rates” language appeared in a FOIA-obtained draft tied to a contract with a reported ceiling of up to $200 million over two years.

OpenAI spokesperson Nate Evans said the company never agreed to that language, characterizing the document as an earlier Pentagon draft that OpenAI rejected. Defense Department spokesperson Jacob Bliss similarly said the phrase did not appear in any active contract.

The factual record, however, remains contested. A Justice Department lawyer initially confirmed the language to The Intercept before the department later walked back that position. Whether the phrase appeared in an executed agreement or only in a rejected draft remains unresolved. It is also unclear whether the language reflected an operational goal even if it was absent from the final text.

Soldiers Can Refuse. Software Cannot, Not in the Same Way.

The legal duty to refuse a manifestly unlawful order is a human obligation, not a technical feature that transfers cleanly to an AI system.

Under the DoD Law of War Manual, service members must refuse clearly unlawful orders, meaning orders that are manifestly or obviously illegal under applicable law. Obedience does not automatically transfer responsibility to the superior who issued the command. The Nuremberg proceedings helped establish that subordinates retain accountability for manifestly unlawful acts, a principle that later shaped modern international criminal law, though the full legal history is more complex than any single formulation captures.

An AI refusal is a different thing entirely. It is a technical control selected by developers, procurement officials, and policymakers, not an exercise of legally recognized individual responsibility. Delegating a decision to software does not eliminate the human duty to make that decision.

The Anthropic comparison sharpens this tension. Anthropic sought contract restrictions on domestic mass surveillance and fully autonomous lethal weapons. The Pentagon required availability for “all lawful purposes” and subsequently excluded Anthropic from classified-network agreements that included, among others, OpenAI, Google, Nvidia, Microsoft, AWS, and SpaceX, according to reporting by The Intercept and Al-Ahram. OpenAI then announced its own Pentagon deal and said it preserved core safeguards, with CEO Sam Altman stating the agreement retained “human responsibility for the use of force, including for autonomous weapon systems,” according to TechCrunch.

Battlefield law rarely arrives as a clean binary. Legal judgments can turn on whether a person is surrendering, wounded, or otherwise protected, and those facts shift faster than any system can reliably process. One design approach worth examining does not require AI to make a final legal determination. Instead, the system flags uncertainty, identifies conflicting signals, and escalates ambiguous cases to a human commander or legal adviser, leaving the human responsible for resolving the issue.

That middle path only works if the accountability chain stays intact. Critics and legal analysts have argued that a model can comply readily and still issue warnings, maintain audit logs, and require human authorization at legally sensitive moments. Reducing refusals without preserving those mechanisms, they contend, does not make a system more capable; it makes accountability harder to trace.

Procurement standards may ultimately need to measure more than accuracy and refusal rates. Documented human authorization, escalation protocols for ambiguous cases, and the ability to suspend a system when legal uncertainty is unresolved are among the criteria advocates have proposed. The duty to say “stop” did not disappear when AI entered the command center. It shifted, and someone still has to own it.

Share this

At Gadget Review, our guides, reviews, and news are driven by thorough human expertise and use our Trust Rating system and the True Score. AI assists in refining our editorial process, ensuring that every article is engaging, clear and succinct. See how we write our content here →