The Army Shipped Agentforce. Anthropic's Models Were Off.

Salesforce cleared the Pentagon's IL5 bar on August 5 by attesting that Anthropic's models were disabled, and Slack was left out at IL4. In an accredited environment the authorization boundary, not your evaluation, decides which models and features you actually run.

By Rajesh Beri·August 6, 2026·12 min read
Share:
A rack-mounted server cabinet inside a windowless, concrete government facility, its heavy steel cage door standing half open, with one horizontal row of blade servers completely dark while every other row glows with sta

Illustration generated using AI

The product you evaluated is not the product you are cleared to deploy. On August 5, 2026, Salesforce announced that Agentforce 360 had cleared Impact Level 5 — and Salesforce officials, answering reporters' questions, indicated to DefenseScoop that the company had to attest to the Pentagon that generative AI models and capabilities supplied by Anthropic were disabled, in order to achieve it. The same announcement notes that IL5 authorization for Agentforce 360 excludes Slack, which sits one tier down at IL4 through GovSlack.

Same vendor. Same product name. Same $5.6 billion contract. A different model roster and a different feature set — decided by an accreditation package, not by anything on your evaluation scorecard. If you are buying an agentic platform into a regulated or accredited environment, that is the finding: the authorization boundary picks your models, and it picks them after you have signed.


What Salesforce Switched Off to Clear IL5

Salesforce had to remove a model provider it had spent a year integrating. In October 2025, Salesforce and Anthropic expanded their partnership to make Claude a preferred model for regulated industries on the Agentforce 360 Platform, with Anthropic described as "the first LLM provider fully integrated within the Salesforce trust boundary" and all Claude traffic contained inside Salesforce's virtual private cloud. Financial services, healthcare, cybersecurity and life sciences were the named targets. CrowdStrike and RBC Wealth Management were the named references.

Ten months later, that integration is the thing that had to come out. Salesforce's position, per DefenseScoop, is that the platform remains model-agnostic and a policy toggle could re-admit Anthropic if the Department of War reverses course. That is a fair description of the architecture. It is also an admission about the operating model: the roster is a runtime setting somebody else controls, and in an accredited enclave that somebody is an authorizing official, not your architecture review board.

The gap is not theoretical, and it is not confined to the Pentagon. Salesforce's own Government Cloud documentation for Agentforce states that the Atlas Reasoning Engine supports only Azure OpenAI GPT-4o in Government Cloud, and that Agentforce for Developers and Agentforce Vibes are not supported there at all. Three environments, three different products: commercial Agentforce with the full model catalogue, FedRAMP High Government Cloud Plus with a reduced one, and the IL5 Missionforce enclave with Anthropic switched off entirely.


Why IL5 Is Exactly Where the Anthropic Ban Bites

The overlap between the authorization and the exclusion is structural, not coincidental. IL5, per DISA's Cloud Computing Security Requirements Guide, covers two things: Controlled Unclassified Information that needs more protection than IL4 affords, and National Security Systems — systems involving intelligence activities, cryptologic activities, command and control of military forces, equipment integral to weapons systems, and functions critical to military or intelligence missions. It is built on a FedRAMP High provisional authorization supplemented with DoD FedRAMP+ controls, requires physical separation from non-federal tenants, and restricts cloud-provider personnel with access to US citizens, nationals or persons.

Now look at the other side. The Anthropic supply-chain risk designation runs on 10 U.S.C. § 3252, which lets the Secretary exclude a company from contracts and subcontracts for National Security Systems. The statute's own scope is the same NSS boundary that defines the upper half of IL5. An IL5 authorization package and a §3252 exclusion therefore overlap by construction — not across all of IL5, which also carries controlled unclassified information the statute does not reach, but across precisely the NSS half.

The sequence that got here is short and documented. Per a timeline compiled by Tech Policy Press, Secretary Pete Hegseth gave Anthropic a February 27, 2026 deadline to permit unrestricted use of its models for all legal purposes; Anthropic declined on February 26; on February 27 President Trump directed federal agencies to cease using Anthropic technology and Hegseth designated the firm a supply-chain risk. The Congressional Research Service records the rest: contractors, suppliers and partners barred from working with Anthropic, an up-to-six-month transition period, a March 9 pair of lawsuits, a March 26 preliminary injunction for Anthropic, and an April 8 appellate ruling that let the designation stand pending review.

That six-month clock, started February 27, runs out this month.


An Accredited Model Can Still Be Switched Off

The most useful distinction in this story is one almost nobody puts in an RFP: a model holding the accreditation for your impact level and a model being permitted in your accredited environment are two different facts, and the second one is decided by policy.

Anthropic's models are in AWS GovCloud right now, served through Amazon Bedrock — the same runtime the commercial Agentforce integration used. Amazon's regional availability matrix lists Claude Sonnet 5 and Claude Opus 4.8 as in-region in GovCloud (US-West), with cross-region access to both from GovCloud (US-East). And they are not merely present: AWS's own compliance listing for Bedrock models carries the Claude family, Opus 4.8 and Sonnet 5 included, as FedRAMP High and DoD CC SRG IL4/IL5 approved. Anthropic states it plainly too — Claude via AWS Bedrock is IL5 accredited, which is why it is the only path it offers for ITAR data.

So the models Salesforce switched off to clear IL5 already hold IL5 accreditation. The accreditation was never the obstacle. A supply-chain designation sits above it, and it can disqualify a model that passed every control an authorizing official would otherwise check.

This is the trap that a proof-of-concept cannot surface. Your team runs the pilot in a commercial tenant against the model you benchmarked, the eval results look good, procurement signs, and then the accredited deployment lands on a model that was never in the test harness. Every eval score, every prompt tuned against a specific model's instruction-following, every latency and cost assumption in the business case — all of it was measured on a system you are not going to run. We wrote about a model swap invalidating an eval suite when the vendor changed the weights underneath a stable version string. This is the same failure with a compliance officer holding the switch instead of a release engineer.


Slack Was in the Contract and Out of the Authorization

The Slack exclusion is the cleaner illustration, because it is a whole product rather than a component. When the Army awarded Salesforce a $5.6 billion, 10-year IDIQ contract in January 2026 — five-year base, five-year option, executed through the Computable Insights LLC subsidiary — the products named in the announcement included Salesforce CRM, Slack and MuleSoft.

Slack is in the contract. Slack is not in the IL5 authorization. GovSlack received its DoD IL4 provisional authorization in December 2025, and IL4 is where it stays. So the collaboration surface where the October 2025 Anthropic partnership put Claude — invoking it directly in Slack to analyze documents and pull Salesforce CRM or Tableau data — is a tier below the agent platform it was meant to sit beside, running a model that is disabled next door.

None of this is a Salesforce failing. Microsoft ran the same pattern: Microsoft 365 Copilot only became available to GCC High customers on December 3, 2025, and shipped with a backlog of additional features promised for the first half of 2026. Accredited environments trail commercial ones by design, because the accreditation is a point-in-time assessment of a fixed configuration and every new capability reopens it. What changed in 2026 is that the thing lagging is no longer a feature — it is the model doing the reasoning.


The Army Bought a Real Thing, Not a Demo

It would be easy to read all this as a compliance curiosity attached to a small deployment. It is not. Army Human Resources Command is the first Department of War organization to deploy the newly IL5-authorized Agentforce, and the scope Salesforce published for it is 9.2 million soldiers, veterans and military families, more than 1,500 cases a day, roughly 600,000 cases a year, over 55 million agent conversations a month at full scale, and $6 million in projected annual savings from reduced manual processing, better case routing and retired legacy systems.

Read the workload description carefully and the design is conservative in a way most enterprise agent programs are not. The agents answer routine inquiries, summarize case histories, and surface policy and career information from approved Army sources — then route complex matters to human specialists for decision-making. Retrieval and summarization over a curated corpus, with a hard handoff at the point of judgment. That is a defensible architecture precisely because it survives a model substitution better than an autonomous one would. An agent that summarizes an approved source degrades gracefully when the model underneath it changes. An agent that decides does not.

Those figures are Salesforce's projections at full scale, not audited results, and the $6 million is a savings estimate rather than a booked number. Treat them as the vendor's case. The architecture is still the lesson.


The Case That This Is Narrower Than It Looks

The strongest counterargument is that the designation's legal reach is far smaller than the headlines imply, and it deserves a fair hearing. Just Security's reading is that §3252 authority is limited to excluding a company from a contract or subcontract for an NSS system. It does not mandate divestment, does not restrict non-NSS uses such as administrative, finance, HR or logistics systems, does not reach a contractor's internal IT or its own commercial contracts, and is not a sanctions authority. The statute explicitly excludes routine administrative and business applications.

By that reading, a defense contractor can keep using Claude for its own engineering and back office, and this is a narrow procurement action about one class of system.

Two things complicate it. First, the executive direction was broader than the statute — the President's order told agencies to cease use, not merely to stop buying for NSS. Second, the compliance machinery is the burden regardless of scope. Under FAR 52.204-30, as Mayer Brown lays out, contractors must check SAM.gov quarterly for new FASCSA orders, conduct a reasonable inquiry into whether covered articles were used in performance, report to the contracting officer within three business days if one was, and file a mitigation plan within ten — and those obligations flow down to all subcontractor tiers. "Covered article" is defined broadly enough to include software and services with embedded or incidental information technology.

That is the part that reaches ordinary enterprises. You do not have to sell to the Department of War to inherit a flow-down clause from a prime who does. And "reasonable inquiry" into which models are embedded in your SaaS stack is not a question most CIOs can answer today, because the vendor's model roster is not in the contract. We covered a single GSA clause reshaping $91.8 billion of federal AI purchasing; this is that mechanism aimed at one company, with a three-business-day reporting trigger attached.


What to Ask Before You Sign the Accredited SKU

This Week: Ask each agentic platform vendor, in writing, for the model roster of the specific environment you will deploy into — commercial, FedRAMP Moderate, FedRAMP High, IL4, IL5 — and the list of features excluded from that authorization. Salesforce publishes this; most vendors will not until asked. Then check whether your last evaluation was run against any model on that list.

This Month: Inventory where a model provider is embedded rather than contracted. Your exposure is not the API key your platform team signed for; it is the model your CRM, your service desk, your security tooling and your document platform call on your behalf. If you hold or subcontract under a federal contract, the six-month transition period that started February 27, 2026 expires this month — the reasonable-inquiry obligation is live now, not at renewal.

Before Renewal: Get three clauses into the contract. One, notification with a defined window before a model is added to or removed from your authorized environment. Two, the right to re-run acceptance evals after a roster change, with a remedy if quality regresses against an agreed baseline. Three, feature-parity disclosure — a maintained list of what the accredited edition does not do, which is the only honest way to price it against the commercial demo you were shown.

Design for it now: Assume the model will change and you will not choose when. Keep prompts and tool definitions portable, keep an eval suite that runs against any provider, and put the human handoff at the point of judgment rather than the point of escalation. The Army's design does exactly this, and it is the reason a model swap is survivable there. Our model-agnostic architecture piece has the longer version of that argument, and Salesforce's own Agent Fabric work is what a multi-vendor control plane looks like when it is designed rather than retrofitted.


The Bottom Line

Enterprise software has always shipped a smaller product into regulated environments — fewer integrations, older builds, a features-coming-soon page. Buyers priced that in. What is new is that the missing component is now the model, and the model is where the capability lives. A CRM missing a dashboard is a CRM. An agent platform running a model you never tested is a different system wearing the same name.

The Pentagon's exclusion of Anthropic is a specific fight with a specific company, and it may not survive the litigation. The mechanism will. Any government, any regulator, any sovereign-cloud authority can now disqualify a model provider and force a substitution inside a platform an enterprise already bought — and we have watched a lab lose federal access with no warning and no failover, watched the Pentagon assemble an eight-vendor roster around one exclusion, and watched that dispute land in federal court. None of those were the buyer's decision either.

You did not choose the model. You chose the boundary. Read it before you sign it.

Continue Reading

Share:

Frequently Asked Questions

What is DoD Impact Level 5 (IL5)?

IL5 is the tier of DISA's Cloud Computing Security Requirements Guide covering Controlled Unclassified Information that needs more protection than IL4 provides, plus unclassified National Security Systems data. It is built on a FedRAMP High provisional authorization supplemented with DoD FedRAMP+ controls, requires physical separation from non-federal tenants, and restricts cloud-provider personnel with access to US citizens, nationals or persons.

Is Slack included in the Agentforce 360 IL5 authorization?

No. Salesforce's announcement states the IL5 authorization for Agentforce 360 excludes Slack, where GovSlack is currently authorized at IL4. Slack was named among the products covered by the $5.6 billion Army IDIQ contract awarded in January 2026, so it is in the contract but a tier below the agent platform in the accreditation.

Are Anthropic's Claude models accredited for DoD IL5?

Yes, in AWS Bedrock. AWS's compliance listing for Bedrock models carries the Claude family — including Claude Opus 4.8 and Claude Sonnet 5 — as FedRAMP High and DoD CC SRG IL4/IL5 approved in AWS GovCloud, and Anthropic points to Bedrock as its IL5-accredited path for ITAR data. That is what makes the Agentforce case instructive rather than routine: the models Salesforce attested were disabled to clear IL5 already hold IL5 accreditation. The exclusion came from the supply-chain risk designation, which sits above the accreditation, not from any security control the models failed.

What does the Army Human Resources Command deployment actually cover?

Per Salesforce, HRC is the first Department of War organization to deploy the newly IL5-authorized Agentforce, serving 9.2 million soldiers, veterans and military families across more than 1,500 cases a day and roughly 600,000 a year, with over 55 million agent conversations a month projected at full scale and $6 million in projected annual savings. The agents answer routine inquiries, summarize case histories and surface policy from approved Army sources, routing complex matters to human specialists. Those figures are vendor projections, not audited results.

Does the Anthropic designation affect companies that do not sell to the Pentagon?

It can reach them through flow-down. Under FAR 52.204-30, contractors must review SAM.gov quarterly for new FASCSA orders, conduct a reasonable inquiry into whether covered articles were used in contract performance, report to the contracting officer within three business days if one was, and file a mitigation plan within ten — and those obligations apply at all subcontractor tiers. Separately, legal analysts note the statutory exclusion authority itself is limited to National Security Systems and does not restrict a contractor's internal IT or purely commercial use.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →