The word missing from OpenAI’s public materials about its Pentagon work is the one that would matter most to anyone worried about how an AI system behaves under military pressure: refusal. The company’s announcement of its agreement with the Department of Defense talks about frontier capabilities, national security missions, and administrative efficiency. It does not talk about what the model will say no to, or how often, or who gets to decide.

That silence is the story.

According to documents obtained by The Intercept through a Freedom of Information Act lawsuit, the Department of Defense asked OpenAI for a custom version of its technology engineered around a specific design goal: minimal refusal rates. The phrase appears in contract materials tied to an expanded prototype deal worth up to $200 million over two years. OpenAI says the language never made it into the executed agreement. The Pentagon’s lawyers first confirmed the document was the final signed contract, then reversed themselves.

Somewhere between those two positions sits the question of what the world’s most valuable AI company has actually agreed to build for the world’s largest military.

Consider how this reads to someone outside the defense procurement world. Contract version numbers, executed versus draft, mod language that survives negotiation and mod language that gets struck — this is the texture of procurement work. The idea that a $200 million agreement would generate confusion at the Pentagon’s own legal office about which document is the final one is unusual. Contracts of that size have very clear paper trails. Someone knows which version was signed.

The dispute over the phrase is, on its face, narrow. Either the words minimal refusal rates appear in a binding document or they do not. But the reason the phrase matters has nothing to do with contract law and everything to do with what an AI safety refusal actually is.

When ChatGPT declines to help someone prioritize drone strike targets, that refusal is not a bug. It is the visible surface of a set of policies, training choices, and reinforcement decisions the company has made about what the system will and will not do. Strip those refusals out and you have not made the model more useful in a neutral sense. You have shifted the moral architecture of the tool from the company’s guardrails to the operator’s discretion.

Heidy Khlaaf, chief scientist at the AI Now Institute and a former systems safety engineer at OpenAI, put the technical reading plainly to The Intercept: “Minimal refusal is likely referring to little or no safeguards on the model being used.” That is the sentence to sit with. Safeguards and refusals are, in machine learning terms, close to the same thing. The refusal is the safeguard rendered as an output.

pentagon building exterior
Photo by Chris on Pexels

OpenAI’s public position, delivered by spokesperson Nate Evans, is that the company never agreed to contract language requiring ‘minimal refusal rates’ and that the phrase does not appear in its executed contract. Evans told The Intercept the released document appears to be an earlier draft proposed by the department before OpenAI gave feedback, that OpenAI rejected the language, and that the department agreed to remove it. The Pentagon, through spokesperson Jacob Bliss, says the phrase does not appear in any active Department of War contract with OpenAI. Both statements can be true while still leaving the underlying question unresolved: what design targets, spoken or unspoken, are shaping the custom military version of the model?

The context matters. In 2025, OpenAI, Google, xAI, and Anthropic all agreed to develop militarized AI prototypes for Pentagon use across logistics, intelligence decision-making, and warfighting. This is not a fringe experiment. It is the mainstream direction of the frontier labs, moving in near-lockstep. It has not been frictionless. Anthropic’s own expansion onto classified networks collapsed earlier this year over whether the contract would bar the use of its technology for autonomous weapons systems and domestic surveillance; the administration then designated the company a supply-chain risk and moved to ban it from government use, a designation a federal judge overturned in late August 2026. The relationship between the labs and the department is close, competitive, and unstable.

Which is what makes the refusal question load-bearing.

The culture around targeting decisions in military intelligence is layered by design. Analysts, JAG officers, commanders. Each layer exists partly to slow the process down enough for someone to say no. An AI system that has been tuned to say no less often does not remove those human layers. But it changes what those humans are pushing back against. Instead of a tool that returns a hard refusal on certain queries, they get a tool that returns a plausible answer, phrased with confidence, that they now have to affirmatively override.

That is a different cognitive task. The default flips.

Refusal rate is one of the most measurable things a model produces. Companies A/B test it, benchmark it, and tune it. When a customer asks for a lower refusal rate, engineers know exactly which levers to pull: adjust the system prompt, retrain on examples where the previous version refused, weaken the reinforcement signal on certain categories. None of this requires a contract clause. It requires an understood expectation. The clause, if it existed, would only be the written form of a preference the buyer could communicate a dozen other ways.

This is where the draft-versus-final debate starts to feel beside the point. If the Pentagon wrote the phrase into a draft, that tells you what the buyer wanted. Whether OpenAI signed a document with those exact words or agreed to the same outcome through less legally exposed channels is a question the FOIA record cannot fully answer.

server room lights
Photo by panumas nikhomkhai on Pexels

Khlaaf’s second observation to The Intercept goes further. She called it a worrying development that private corporations are given the power to make determinations in warfare that are ultimately state obligations — whether those are targeting recommendations or the guardrails deployed to constrain a state’s military use of the technology. The framing is worth pausing on. In the traditional model, the state sets the rules of engagement and the tools execute them. When the tool itself embeds judgment about what queries to answer and what to refuse, the company selling the tool has effectively taken on a piece of the rules-of-engagement function. Loosening those refusals is not deregulation in a neutral sense. It is a transfer of decision authority from the vendor’s policy team to the operator’s chain of command.

Whether that transfer is appropriate is a legitimate debate. It is also a debate that has largely not happened in public, because the terms of these contracts are negotiated inside classified and semi-classified channels, surfaced only when reporters like The Intercept pry documents loose. The deployment contract materials now in the public record are unusual precisely because they are in the public record.

There is a wider pattern here that connects to how quickly powerful tools are being handed to government users without the procedural constraints that would normally accompany them. The speed of AI procurement inside the national security apparatus has outrun the institutional muscle memory for interrogating what these systems will and will not do. The contracts are being written before the norms are.

In commercial contracting, an ambiguous artifact of this size would trigger an immediate audit. In defense contracting involving a rapidly moving technology, ambiguity is sometimes the point. It gives both parties room to shape the deliverable through the working relationship rather than the paper.

So the honest thing to say about the current state of the record is this. A phrase that would represent a meaningful weakening of AI safety architecture appeared in a Pentagon contract document. OpenAI says it did not survive into the final agreement. The Pentagon initially said it did, then said it didn’t. The underlying capability the phrase describes — a model tuned to refuse military commands less often — is technically straightforward to build regardless of whether the phrase appears anywhere in writing.

The public has been given a dispute about words. The engineering question sits underneath it, unanswered.

What OpenAI will not say plainly, and what the Pentagon will not confirm, is the number that would settle this. How often does the military version refuse a query the consumer version would refuse? If that gap is small, the safeguards remain in place. If it is large, they do not. Everything else is contract theater around a measurement the buyer and seller already have.