If your rule says prompts must be processed in region, buy the model through a hyperscaler, not from the lab that built it. Amazon Bedrock geographic inference profiles and Microsoft Foundry Data Zone deployments keep processing, stored data and abuse-monitoring data inside the boundary, and a single policy can stop global routing. OpenAI's own API processes prompts in region only in the US, the EU and the UAE. Anthropic's own API does so only in the US.
Every one of these paths charges roughly the same 10% for the privilege, so the premium should not decide this. What decides it is what each residency promise leaves out. Prices below were checked on 15 September 2026 against each vendor's live pricing page or price API.
| Path | Where prompts are processed in region | Main caveat | In-region premium | Enforcement lever |
|---|---|---|---|---|
| Amazon Bedrock, geographic profile | Geographies such as US, EU, APAC | Abuse-detection copies for some GPT and Fable models | Global profiles ~10% cheaper | SCP can deny global profiles |
| Microsoft Foundry, Data Zone or Standard | US, EU or APAC zone, or one Azure geography | Abuse-monitoring store and human review unless exempted | GPT-5.5: $5.50/$33 vs $5/$30 | Azure Policy can deny Global SKUs |
| Google Vertex AI, regional or multi-region endpoint | Regional endpoints; US and EU multi-region | The global endpoint carries no guarantee | Global "generally lower"; Claude +10% | Org policy can restrict endpoints |
| OpenAI API, regional project | US, EU (EEA + Switzerland), UAE | Seven more regions are storage-only | +10% on models released from 5 Mar 2026 | Project geography |
| Mistral AI, EU or US endpoint | EU, US | Control-plane metadata; agents, batch, files | 1.1x | Endpoint hostname per client |
| Anthropic Claude API | US only | Storage is US-only; flagged content may be kept up to 2 years | 1.1x | Workspace geo allowlist |
Why Bedrock Is the Default Pick for a Hard Residency Rule
Bedrock wins because its defaults are the ones a compliance team would have written itself. Amazon Bedrock runs a zero operator access model and by default does not store model inputs or outputs. A geographic cross-Region inference profile keeps a request inside its geography: a US request stays in US Regions. The same page says plainly that "your input prompts and output results might move outside of your source Region" within that geography.
Read that sentence twice. The guarantee is in-geography, not in-Region. If your rule names one country inside the EU, a geographic profile is the wrong control. Before you assume there is a right one, check whether your model can be invoked in a single Region at all.
The enforcement is what makes it the pick. AWS documents a service control policy that denies global cross-Region inference by matching aws:RequestedRegion set to unspecified on global.* inference profiles. That is one policy at the organization root. Below it, no developer can move a workload onto worldwide routing by editing a model ID.
The exceptions are narrow, and AWS writes them down. For the OpenAI GPT models Bedrock hosts, classifier-flagged traffic is kept for up to 30 days for offline abuse detection. For Claude Fable 5 and Fable 5.1, all traffic is kept for up to 30 days, and flagged traffic may be reviewed by AWS staff. With cross-Region inference enabled, those copies are stored in the destination Region. They stay in your geography, not necessarily in your source Region.
The steel-man for Microsoft Foundry instead: if your estate already runs on GPT models and Entra ID, and you have a Microsoft account team, Data Zone gives you a residency story that is nearly as complete. The next section shows where it falls short.
Which LLM Providers Can Process Prompts In Region?
Region coverage is uneven, and the region count on most vendor pages counts storage, not processing. Here is what each provider's own documentation actually commits to.
- OpenAI API. OpenAI's data controls page lists ten regions, but only three support regional processing: the United States, Europe (EEA plus Switzerland) and the United Arab Emirates. Australia, Canada, India, Japan, Singapore, South Korea and the United Kingdom offer storage at rest only. Any region other than the US also requires approval for abuse-monitoring controls and a signed Modified Retention amendment.
- Anthropic Claude API. The
inference_geoparameter accepts only"us"or"global", and workspace geo, which controls where data is stored at rest, is currently"us"only. The parameter works on Claude 4.6 and later models. Sending it to older models returns a 400 error. - Microsoft Foundry (Azure OpenAI). Data Zone deployments process prompts within the US, EU or APAC zone, and Standard deployments within one Azure geography. The EU zone follows the EU Data Boundary, which "can include" Norway and Switzerland. Microsoft "can add regions to either data zone without prior notice."
- Amazon Bedrock. Geographic profiles cover geographies such as the US, EU and APAC, as described above. Bedrock sets the price by the source Region you call from.
- Google Vertex AI. Google Vertex AI offers regional endpoints plus US and EU multi-region endpoints for Claude, generally available since 15 May 2026. Google says these ensure "your data processing remains within your preferred geography."
- Mistral AI. Regional endpoints at
api.eu.mistral.aiandapi.us.mistral.aibecame generally available on 11 August 2026.
The practical consequence is blunt. A UK or Japanese buyer whose rule covers processing cannot meet it on OpenAI's or Anthropic's own APIs at all. That buyer should confirm that the specific model is available in the specific hyperscaler geography before assuming anyone can meet it.
Storage Is Not Processing, and Your Contract Probably Blurs Them
Data residency, strictly, is where your content is stored at rest. In-region processing is where the model actually computes over it. Vendors often sell the first under the name of the second, and questionnaires rarely force the distinction.
OpenAI's storage-only regions are the sharpest case. Its own documentation warns that "extended prompt caching in regions that do not support Regional processing may require that OpenAI process and temporarily store Customer Content outside of the Region." A project pinned to the UK can store data in the UK and still compute on it somewhere else.
Mistral draws the line in a different place. Its regional endpoints keep inference in region, but "account configuration, API keys, billing, access management, usage analytics, and other operational metadata may still be handled by Mistral systems outside the selected inference geography," per Mistral's regional inference docs. If your policy counts usage logs as regulated data, that is a finding, not a footnote. (We covered the commercial side of that endpoint in Mistral Wants Multi-Year Money. Its EU Endpoint Drops Agents.)
Anthropic runs the other way round: inference can be pinned, but only to the US, and storage has no non-US option at all.
What Data Residency Silently Excludes
Every residency promise comes with a feature list. The features left off it are the ones teams tend to add in month three: batch jobs, caching, agents and tools that call out.
- Batch. On Foundry, Global Batch processes "in any Azure region"; Data Zone Batch stays in the zone. Both carry the same 50% discount, so a cost-driven migration to batch can quietly pick the wrong SKU. Mistral's regional endpoints do not support batch at all.
- Agents, files and tools. Mistral's regional endpoints support chat completions, and "function calling is the only supported regional tool." Agents, Batch and the Files API are not available regionally.
- Remote tools. OpenAI notes that MCP servers used with its remote MCP tool "are third-party services" whose own data residency policies apply. Residency on the model call says nothing about the tool call.
- The wrong Google product. Google's Gemini Developer API ZDR page says that with Grounding with Google Search, Google stores prompts and output for 30 days and there is "no way to disable" it. It also says: "If your workload requires guaranteed zero data retention or enterprise data processing agreements, use Vertex AI." A prototype built on a developer API key is outside the enterprise terms you negotiated.
- Retention you cannot contract away. Anthropic may keep inputs and outputs for up to 2 years if a session is flagged by its trust and safety systems, even under ZDR. Its Fable 5, Fable 5.1, Mythos 5 and Mythos 5.1 models require 30-day retention unless Anthropic expressly authorizes otherwise. (More in Fable 5 Needs Retention On. ZDR Was Never Zero.)
Abuse Monitoring Is the Exception That Matters Most
Abuse monitoring is the one process that stores your prompts on purpose. Where it stores them, and who may read them, is the real residency test.
Microsoft Foundry comes closest to answering it in writing. For Global and Data Zone deployments, the abuse-monitoring data store sits in the customer's designated geography. Human reviewers see prompts only when they have been flagged or form part of an abusive pattern. They work from Secure Access Workstations with just-in-time approval, and "for Models sold by Azure deployed in the European Economic Area, the authorized Microsoft employees are located in the European Economic Area."
Turning storage and human review off is a separate matter. Modified abuse monitoring requires meeting Limited Access eligibility criteria, and even when granted, automated review may still run. In an August 2026 answer to a nonprofit asking about EU residency, a Microsoft external-staff moderator said the exemption is "available only to customers and partners managed by a Microsoft account team or under an eligible program." The same answer flagged the more basic trap: a Global Standard deployment in an EU region can process prompts outside the EU.
OpenAI makes the exemption a precondition, not an option. Abuse-monitoring logs are kept for up to 30 days by default, and outside the US you must already hold approved abuse-monitoring controls before residency is available. Getting approved is its own project. In an October 2025 thread on OpenAI's developer forum, a Tier 5 API customer described contacting sales twice for data residency and getting back a no-reply message recommending "self-serve resources."
Bedrock stores nothing by default, apart from the named-model exceptions above.
Who Is Actually Processing Your Prompts?
The path you buy through decides who the data processor is, and the processor decides the sub-processor chain behind it.
- Foundry: Microsoft states that models sold by Azure "do NOT interact with any services operated by providers" of those models, OpenAI included. GPT on Foundry is Microsoft processing, not OpenAI.
- Bedrock: retained abuse-detection data is "not shared with third-party model providers".
- Claude through a cloud: Anthropic says that on Bedrock and Google Cloud, "the cloud provider is the data processor", so its own retention page is not the document that governs you there.
- Mistral: inference stays in the selected region "subject to limited, safeguarded transfers to sub-processors that may occur outside that region," per Mistral's announcement. Read the Trust Center list it points to before you sign.
Put a gateway in front of any of these and you have added a processor whose data policy can change owner. That is exactly what happened in Stripe Bought OpenRouter. A Toggle Is Not a Contract.
Contract Promises vs Technical Controls
A contract tells you what the vendor promises. A technical control tells you what your own developers cannot undo. You need both, and most buyers stop at the first.
Controls that hold without trusting every developer:
- AWS: the SCP above, set at the organization root.
- Azure: an Azure Policy that denies the Global SKUs:
GlobalStandard,GlobalProvisionedManagedandGlobalBatch, plusDeveloperTier, which also processes in any Azure region. DenyingGlobalStandardalone leaves the other three open. After an exemption is granted, verify it with the resource'sContentLoggingcapability, which readsfalseonly when abuse-monitoring storage is off. - Anthropic: set
allowed_inference_geosto["us"]on the workspace, so any request asking for another geo errors out. Then check that theusage.inference_geofield in each response says where inference actually ran.
Controls that depend on every client getting it right:
- OpenAI lets a key from a project with Global geography opt into regional processing per request by calling the prefixed domain. The same flexibility means any client that skips the prefix is processed globally. Use dedicated regional projects instead.
- Mistral selects the region with a
serverparameter on each client. The defaultapi.mistral.aiendpoint is global. - Google decides the region by the endpoint the client calls, and Google's own tooling has got that wrong. In a gemini-cli issue, a user with
GOOGLE_CLOUD_LOCATION=europe-west3found that authenticating with an API key sent requests to the global endpoint, while the tool's/aboutscreen still displayed the European region. The issue was closed as not planned. Google's Restrict Endpoint Usage organization policy can deny the globalaiplatform.googleapis.comendpoint, though Google notes such policies "are not data residency commitments." A separate question on Google's developer forum, asking whether pay-as-you-go guarantees in-region processing in Japan, had no visible answer from Google staff when checked on 15 September 2026.
What Does In-Region Inference Actually Cost?
In-region inference costs about 10% more on every path, which on a realistic workload is tens to low hundreds of dollars a month. Take a fixed workload of 100 million input and 20 million output tokens a month, with no caching and no batch discount:
| Model and path | Global list price | In-region price | Monthly premium |
|---|---|---|---|
Claude Sonnet 5, Claude API, inference_geo: "us" |
$400 | $440 | $40 |
| GPT-5.5, Microsoft Foundry, Data Zone (US East 2) | $1,100 | $1,210 | $110 |
| GPT-5.6 Sol, OpenAI API, regional project | $800 | $880 where the uplift applies | $80 |
The inputs: Claude Sonnet 5 lists at $2/$10 per million tokens, and US-only inference is 1.1x across every token category. Azure's retail price API lists GPT-5.5 at $5.00/$30.00 Global and $5.50/$33.00 Data Zone in US East 2. OpenAI lists GPT-5.6 Sol at $4/$20 and charges a 10% uplift on residency-eligible models released on or after 5 March 2026.
The other paths land in the same place. Mistral bills regional inference at 1.1x list. Bedrock's global profiles save "approximately 10%" against geographic ones. For Claude on Bedrock and Google Cloud, regional and multi-region endpoints carry a 10% premium over global, and Google says Claude prices "are generally lower on global endpoints."
The real cost is model lag, not price. Microsoft states that geography-based deployment types "arrive last, have no guaranteed availability date, and depend on capacity that frees up as older models retire." Anthropic's rate limits are shared across geos. A strict single-geography rule can cost you months on the newest model, and nothing on the price list shows that.
Who Should Not Buy Each Option
Amazon Bedrock: skip it if your rule names a single Region and your model is only served through a multi-Region inference profile. Also skip it if you plan to run Claude Fable 5 under ZDR past 31 December 2026, the date Bedrock's Enterprise Frontier Safeguards ZDR ends and 30-day retention begins.
Microsoft Foundry: skip it if you need a new model on launch day in a single-geography deployment. Skip it if you are not a managed Microsoft customer and your policy forbids human review of prompts. And skip it if "EU" in your contract means EU member states only, since the EU zone can include EFTA countries and gain regions without notice.
Google Vertex AI: skip it if your developers use API-key authentication or pick endpoints freely, and you have no way to catch calls to the global endpoint. Claude on Vertex's EU multi-region endpoint is a good Claude-in-Europe answer only if you can enforce the endpoint.
OpenAI API: skip it if you are in the UK, Japan, India, Canada, Australia, Singapore or South Korea and your rule covers processing. Skip it if you have no approved abuse-monitoring controls yet. And skip it if your agents call remote MCP servers with regulated data.
Mistral AI: skip it if you need agents, batch or file storage in region, or if your rule covers operational metadata. Pick it when jurisdiction, not just location, is the requirement. Then accept the feature cut, and check that Mistral Large or the model you need is actually served regionally.
Anthropic Claude API: skip it if you are outside the United States and your rule requires processing in your own region. The next section explains why.
The Loser: Anthropic's Own API, for Anyone Outside the US
Give Anthropic its due first. Its controls are the cleanest design here: a per-request parameter, a workspace allowlist that fails closed, and a response field that tells you where inference ran.
It still loses this comparison, because a well-designed control over one region is not a residency offering. Only "us" and "global" exist for inference, and only "us" for storage. A Frankfurt insurer or a Tokyo bank cannot use it to keep processing in its own region.
Claude itself is not the loser. Buy it through a Bedrock EU inference profile or Google's EU multi-region endpoint, where the cloud provider is the processor and the geography options exist. Check the specific model first: on Bedrock, Claude Fable 5.1's regional endpoints are currently in us-east-1 only. You pay the same 10% for a residency answer that holds up.
Five Questions That Predict Regret
- Does your rule say "stored" or "processed"? If processed, rule out storage-only regions before any demo.
- Which boundary, exactly? A country, the EU, the EEA plus Switzerland, and Microsoft's EU Data Boundary are four different lines. OpenAI's Europe region is the EEA plus Switzerland, and Azure's EU zone can include Norway.
- What is in the request path besides the model? Batch, caching, agents, file storage, grounding, MCP tools and gateways each carry their own residency terms, or none.
- Who stores prompts for abuse monitoring, where, for how long, and who can read them? Get the answer in writing under the data processing addendum, not from a sales deck.
- Can one developer undo it? If residency depends on a hostname or a request parameter, assume someone will eventually leave it out.
What to Do Before Your Next Architecture Review
This Week:
- Ask legal for the exact clause and mark it storage, processing or jurisdiction. If it is jurisdiction, residency is not the answer. Read 7 OpenAI Alternatives. Only 3 Clear a Sovereignty Rule. first.
- List every LLM call path in production, including batch jobs, agents and whatever your developers prototyped on personal API keys.
This Month:
- Enforce below the application: the Bedrock SCP, the Azure Policy SKU deny, or the Anthropic workspace allowlist.
- On Anthropic, log
usage.inference_geoon every response and alert on anything that is notus. - File the modified abuse monitoring or ZDR application now. Treat approval lead time as a project dependency, not paperwork.
Before Renewal:
- Put the sub-processor list, the abuse-monitoring storage location, the reviewer location and any excluded features into the order form as named exhibits.
- Negotiate notice of any change to data zone membership, because Microsoft's documentation reserves the right to add regions without it.
The Bottom Line
Residency is the most mature compliance feature the LLM market sells, and still the easiest to buy wrong. The prices converged on 10%. The promises did not: one covers storage but not processing, one covers inference but not metadata, one covers the model but not the batch job, and one covers a single country. Cloud platforms won this round because they already enforce boundaries for everything else you run, and residency is just one more boundary.
Ten percent buys you a region. The boundary you have to build yourself.
Continue Reading
- 7 OpenAI Alternatives. Only 3 Clear a Sovereignty Rule.
- Fable 5 Needs Retention On. ZDR Was Never Zero.
- Mistral Wants Multi-Year Money. Its EU Endpoint Drops Agents.
- Claude vs GPT vs Gemini: Stop Comparing Per-Token Prices
- Samsung Led Mistral's Round. Score the Clauses, Not the Flag.
- Stripe Bought OpenRouter. A Toggle Is Not a Contract.
