An AI vendor security review comes down to six answers, and a SOC 2 report gives you none of them. Ask these instead:
- Who else processes the prompt, and how much warning do you get when that changes?
- Is "no training" a contract term on every tier your staff actually use?
- What is the longest any copy can live, flagged content included?
- What voids the IP indemnity?
- What testing covers the feature you are deploying, not just the model?
- Is that feature named in the certification scope?
Put those six to the five platforms most enterprises buy directly, using the terms each vendor published as of September 12, 2026. Microsoft gives you the most evidence you can check yourself. Amazon Bedrock gives you the cleanest data boundary. Anthropic's standard paper needs the most negotiation. Anthropic's addendum promises only "reasonable notice" of new subprocessors and a 15-day objection window. Its API docs also say flagged content can be kept for up to two years even under zero data retention.
| The question | OpenAI (ChatGPT Enterprise, API) | Anthropic (Claude Enterprise, API) | Microsoft (Copilot, Foundry) | Google (Gemini in Workspace, Vertex AI) | Amazon Bedrock |
|---|---|---|---|---|---|
| Subprocessor change: notice and exit | 30 days to object; either side may terminate if unresolved | "Reasonable notice"; 15 days to object; remedy is good-faith talks | 6 months, but 30 days for AI subprocessors since May 22, 2026 | At least 30 days; 90 days to terminate for convenience | Page updated at least 30 days ahead; email alerts |
| Trains on business data? | No, by default | No, by contract | No foundation-model training | No, without permission | No, and not shared with model providers |
| Longest documented retention | 30-day API abuse logs; CSAM-scanned images kept even under ZDR | Flagged content up to 2 years, even under ZDR | Flagged samples held for human review; no duration published | Deletion within a maximum of 180 days | Nothing stored by default; 30 days for named frontier models |
| IP indemnity | Copyright Shield: Enterprise and API | IP claims, outside the liability cap | Copyright; Azure OpenAI requires your own evaluation report | Training data and generated output | Uncapped copyright cover, Amazon's own models only |
| AI feature named in ISO 42001 scope? | Certificate listed; product scope not itemised publicly | Commercial products, not consumer plans | Copilot, Copilot Studio, Foundry named | Listed for Workspace | Bedrock named |
| Verdict | Sound default; watch stateful endpoints | Negotiate hardest | Most checkable | Slowest deletion commitment | Cleanest boundary |
The reference case throughout is a 5,000-employee company with staff in the US and the EU. It is approving two things: an assistant for knowledge workers that reads mail and documents, and API access for an internal retrieval app over customer contracts. No health data is involved, and it buys direct from the platform rather than through a reseller. Change any of those facts and the "What Changes the Answer" section below tells you which cell moves.
Why Your Questionnaire Misses the Model Provider
A standard security questionnaire misses the AI supply chain because it was written for software that stores your data, not software that ships it to a third party to think about. DeepInspect's August 2026 review of AI vendor questionnaires put it bluntly: "Most AI vendor security questionnaires are SOC 2 templates with two AI questions tacked on."
The structural reason is how SOC 2 handles dependencies. A subservice organization is a third party a vendor relies on to deliver its service. Under the carve-out method, the vendor's report describes what that third party does but leaves its controls out of the auditor's scope. For an AI feature, the carved-out party is very often the model provider. The clean opinion covers the vendor's side of the line, and the prompt crosses to the other side.
So a SaaS vendor's SOC 2 can be flawless while the thing you are actually worried about — where the prompt goes, how long it stays, who can read it — sits in someone else's contract. That is why the six questions below go to the platform directly. They also go to every vendor that wraps one.
Who Else Touches the Prompt, and How Much Warning Do You Get?
The answer to question one changed for Microsoft customers twice in 2026, which is the best argument for asking it every renewal rather than once. A subprocessor is any third party a vendor engages to process your data on its behalf. A model provider is the subprocessor whose model actually reads the prompt.
Microsoft 365 Copilot — now renamed Microsoft Copilot — onboarded Anthropic as a subprocessor and enables Anthropic models "on by default for most customers in commercial cloud (excluding EU/EFTA and UK)." Registora dates that switch to January 7, 2026. Microsoft also says those Anthropic models "are currently excluded from the EU Data Boundary."
Then OpenAI joined Microsoft's subprocessor list on June 23, 2026. As of July 24 — 31 days later — OpenAI-operated models "are enabled for all users for eligible commercial customers, unless you specifically disable" them. One month earlier, Microsoft's May 22, 2026 Data Protection Addendum update had cut notice for AI subprocessors to 30 days, while other subprocessors keep six months. You retain the right to disable the new AI vendor until at least six months after notice.
None of that is hidden. Every step was documented and every one has an admin toggle. But if your security review cleared Copilot in 2025, the list of labs reading your staff's prompts has since grown by two. And if you are yourself a processor who promised your own customers six months' notice, Microsoft's new window is shorter than yours.
The other four published terms, side by side:
- OpenAI discloses 24 subprocessors. Its DPA gives customers 30 days to object and lets either party terminate the affected services if the objection is not resolved, with prepaid fees refunded.
- Google commits to notify at least 30 days before a new subprocessor starts processing, naming its location and activities, and lets you terminate for convenience within 90 days.
- AWS says it will update its sub-processor page at least 30 days before engaging a new one and emails subscribers.
- Anthropic's Data Processing Addendum promises only "reasonable notice" before a new subprocessor gets access and gives you "fifteen (15) days" to object. The remedy it names is that the parties "will work together in good faith." That is the weakest change-control language of the five.
A subprocessor list also tells you who touches your data, not whose model weights are inside the product. We found that the hard way when Thomson Reuters' legal model turned out to be a Qwen derivative that no contract disclosed. For any vendor that wraps a model, add base model, publisher and licence to the question.
Is "No Training" a Contract Term or a Setting?
All five platforms promise not to train on business-tier data, so the question that actually changes the answer is whether every tier your people use is a business tier. On paper the five match. Anthropic's commercial terms say it "may not train models on Customer Content from Services." OpenAI says API data "is not used to train or improve OpenAI models" and that ChatGPT Enterprise and Business content is not used by default. Microsoft Copilot prompts "aren't used to train foundation models." Google Workspace does not train on customer data "without customer's prior permission or instruction." Bedrock content "is not shared with any model providers."
The gap is the tier your employees choose when they sign up with a personal email. Anthropic's August 28, 2025 consumer-terms change asked Claude Free, Pro and Max users to decide whether their chats improve Claude, extending retention to five years for those who allow it. It explicitly excluded Claude for Work and API use. The contract you negotiated does not follow an employee to a personal plan. That is how one employee's AI tool use became an SEC filing.
Then check how the opt-out actually works, because the mechanism is itself evidence. Slack's privacy principles still say that to exclude your data from its global machine-learning models, a workspace owner must email feedback@slack.com with the subject line "Slack Global model opt-out request." That line drew a public backlash in May 2024. Slack does say it will not train generative models on customer data without opt-in. An opt-out that lives in an inbox is still one you cannot audit.
A guarantee can also be scoped to a route rather than a customer. Cursor's zero data retention stops applying when a developer brings their own API key — same tool, same approval, different data path.
What Is the Longest Any Copy Can Live, and What Proves Deletion?
The retention number that matters is the ceiling, not the default, and the vendors who publish an honest ceiling look worse on paper than those who publish none. Zero data retention (ZDR) is an arrangement under which the vendor does not store prompts and outputs after the request completes — subject to exceptions every vendor documents.
Anthropic's ceiling is the longest, and it is stated plainly in its API documentation: "Even with ZDR or HIPAA arrangements in place… if a chat or session is flagged, Anthropic may retain inputs and outputs for up to 2 years." Its commercial retention policy adds that trust-and-safety classification scores are kept for up to seven years. Its Covered Models — Claude Fable 5 and 5.1, Mythos 5 and 5.1 — require 30-day retention, so ZDR is unavailable for them unless Anthropic expressly authorizes it, which we covered in August.
OpenAI's default is shorter, and its exceptions sit on specific endpoints. API abuse-monitoring logs are kept for up to 30 days. ZDR requires OpenAI's approval, and /v1/assistants, /v1/threads, /v1/vector_stores and /v1/conversations store data until you delete it. Image and file inputs scanned for CSAM are "retained for manual review, even if Zero Data Retention… is enabled."
Retention promises also bend to courts. In the New York Times copyright case, a federal magistrate judge ordered OpenAI to preserve output logs that would otherwise have been deleted. That obligation ended on September 26, 2025, though logs already saved and accounts the Times flagged stay preserved. ChatGPT Enterprise was outside the order's scope. That exclusion is the practical value of buying the business tier: a legal hold aimed at consumer data never reached it.
Microsoft wins this question on evidence. Foundry's abuse monitoring holds flagged prompts for human review without publishing a duration, which is a gap worth raising in writing. But customers approved for modified abuse monitoring can verify it themselves: a ContentLogging capability set to false appears on the resource only when that storage is off. That is a deletion artifact you can screenshot for an auditor, rather than a promise.
The trap on Copilot is web grounding. Queries sent to Bing are not covered by the DPA — Microsoft acts as an independent data controller for them. The EU Data Boundary and HIPAA coverage do not apply to those queries either.
Google's commitment is the slowest. Its data processing addendum deletes data you delete "within a maximum period of 180 days." Separately, Gemini app conversations in Workspace auto-delete after 18 months by default, adjustable by admins to 3 or 36 months.
Bedrock has the cleanest default. It stores no inputs or outputs by default and runs zero operator access, and model providers cannot reach the accounts that host their models. The exceptions are named. Flagged traffic on OpenAI's GPT-5.x and GPT-6 Astra models is kept for up to 30 days. For Claude Fable 5 and 5.1, "all traffic will be retained for up to 30 days," with flagged traffic open to human review by AWS.
What Voids the IP Indemnity?
Every vendor's IP indemnity is conditional, so the question is which condition your own deployment is most likely to trip. An IP indemnity is the vendor's promise to defend you, and pay, if a third party sues over its service or its output.
- Anthropic covers claims that authorized paid use or outputs violate "any third-party intellectual property right," and that obligation sits outside the 12-month-fees liability cap. It carves out modified outputs, combination with other technology, customer inputs, and use the customer "knows or reasonably should know" infringes. This is the strongest indemnity term in the group, and it is why "negotiate hardest" is not "avoid".
- OpenAI's Copyright Shield covers ChatGPT Enterprise and the API, not free ChatGPT or Plus, and is limited to copyright.
- Microsoft's Customer Copyright Commitment extends to Anthropic and OpenAI models inside covered products like Copilot. The exception is "Anthropic models with Data Retention", which run under Anthropic's own terms. On Azure OpenAI the commitment requires a metaprompt and a customer-run evaluation whose report "must be retained by the customer and provided to Microsoft in the event of a claim."
- Google's indemnity has two prongs, training data and generated output, and applies only if you "didn't try to intentionally create or use generated output to infringe."
- AWS offers uncapped indemnity against copyright claims on outputs of its "Indemnified Generative AI Services" — Amazon's own models and services listed in its Service Terms. Third-party models on Bedrock may carry their provider's terms, and Anthropic's models there are governed by Anthropic's Commercial Terms.
That last point is Bedrock's biggest trap. If the indemnity is why you chose Bedrock, check whose model you are actually calling. The circumvention clauses are the other common trap, and a re-prompting employee can trip them.
Whose Red-Team Evidence Covers Your Deployment?
Every major lab publishes model-level safety evaluations, and none of it tests the deployment you are about to approve. The shelf is full:
- OpenAI's Deployment Safety Hub posted the GPT-6 Astra system card on September 3, 2026.
- Anthropic's system cards cover Claude Fable 5.1 and Mythos 5.1.
- Google DeepMind's model cards cover Gemini 3.8 Flash.
- Microsoft's Copilot application card lists risk and safety evaluations, including direct and indirect jailbreak.
The best evidence that this is not enough is EchoLeak. In June 2025, Aim Labs disclosed CVE-2025-32711, a zero-click exploit in Microsoft 365 Copilot. A single crafted email made Copilot exfiltrate internal data by chaining four bypasses:
- past Microsoft's cross-prompt-injection classifier;
- around link redaction, using reference-style Markdown;
- through auto-fetched images;
- via a Teams endpoint the content security policy allowed.
Microsoft fixed it server-side in May 2025, before disclosure, and found no exploitation in the wild. Every model card was accurate. The deployment was still exploitable.
So ask for application-level evidence against the risks that belong to your use case. For an assistant that reads mail, that is OWASP's LLM01:2025 Prompt Injection; for an agent that takes actions, LLM06 Excessive Agency. Even AWS's own Nova Act service card says "we cannot guarantee all prompt injections attacks will be deflected." It also reports blocking 96.4% of harmful prompts on a proprietary dataset — a vendor figure on the vendor's test set.
Microsoft is the only one of the five that turns this into an obligation on you. It makes your own testing report a condition of its copyright commitment on Azure OpenAI, and that is the right instinct. Treat a vendor's test result as a starting point: a benchmark gap is often noise until you size the eval set, and inherited assurance saturates.
Does the Certificate Actually Name the AI Feature?
A certification only helps if the AI feature is named inside its scope, and a certification of the feature still does not test the feature's behaviour. ISO/IEC 42001 is the international standard for an AI management system — the policies, risk processes and oversight an organisation runs around AI. It does not certify that a product is safe.
Scope is where the five differ:
- Microsoft lists GitHub Copilot, Microsoft Copilot, Copilot Chat, Copilot Studio, Foundry and Security Copilot in its ISO 42001 scope. It adds that you remain "responsible, however, for engaging an assessor to evaluate the controls and processes within your own organization."
- AWS's November 2024 certification names Amazon Bedrock, Q Business, Textract and Transcribe.
- Anthropic was certified in January 2025 by Schellman. Its certifications apply to Claude for Work and the API, not Free, Pro or Max.
- OpenAI's trust portal scopes its SOC 2 Type 2 to the API Platform, ChatGPT Enterprise, Edu and Team, for July 1, 2025 to June 30, 2026. It lists ISO 42001 without itemising product scope publicly.
- Google lists ISO/IEC 42001 for Gemini in Workspace.
The sharpest example of scope moving under you is inside one Microsoft product. OpenAI-operated models in Copilot are not FedRAMP High authorized, and no SOC 1 Type 2 report, PCI DSS attestation or HITRUST letter is available for them. The same Copilot, answering through a different model route, carries a different set of attestations.
Microsoft lists those exclusions for the OpenAI-operated route specifically. The product did not change; the path the prompt takes did.
Who Should Not Pick Each Platform
Saying who should walk away is the most useful thing a buyer's guide can do, so here it is for each platform on the reference case.
- Do not pick ChatGPT Enterprise or the API if you need a ChatGPT workspace to keep conversations for less than 90 days — admins cannot set retention below that floor. The same goes if your RAG app depends on Assistants, threads or vector stores and you were counting on ZDR for them.
- Do not pick Claude direct on standard paper if your DPA process requires a fixed notice period and a termination right on subprocessor objections. The same applies if a two-year retention tail on flagged content breaks your records schedule. Both are negotiable for large buyers; neither is the default.
- Do not pick Microsoft Copilot or Foundry without re-reviewing if you promised your own customers six months' notice of AI subprocessors, or if the EU Data Boundary is a hard line and your users will select Claude.
- Do not pick Gemini or Vertex AI if your retention schedule needs vendor-side deletion evidence inside 30 days. The published maximum is 180.
- Do not pick Bedrock for the indemnity if the model you want is Anthropic's or another third party's. The uncapped cover is for Amazon's models.
The loser on this reference case is Anthropic's standard DPA, and it loses on change control rather than on security. It has the strongest indemnity, and it still gives a buyer the least contractual leverage when something changes. For a platform whose models now sit inside Microsoft, AWS and Google as well, that is the clause most worth fixing before signature.
The Case for Trusting the Certificates Instead
The strongest objection to all of this is that it duplicates work auditors are paid to do, and for most SaaS vendors that objection is correct. A SOC 2 Type 2 with a clean opinion, plus ISO 27001, covers access control, change management and incident response better than any questionnaire. CSA's STAR for AI Level 2 now combines a validated AI questionnaire with ISO 42001, which is exactly the standardisation that should shrink these reviews over time.
But the six questions are not about whether the vendor runs good controls. They are about terms and routes, which no attestation covers:
- who the vendor will add next;
- what the carve-out left out;
- which model the request actually reached;
- which condition voids the indemnity.
Keep leaning on the certificates for everything else. Ask the six in writing.
What Changes the Answer
Five facts about your situation move the verdict, and they predict regret better than any feature list.
- You are a processor for your own customers. Subprocessor notice becomes the top criterion, because a vendor's 30-day window has to fit inside your own promise. Google's 90-day termination-for-convenience right and OpenAI's termination right beat a good-faith discussion.
- EU residency is contractual, not aspirational. Model routes decide it. Claude inside Copilot is outside the EU Data Boundary today, and Bing web grounding is outside the DPA entirely.
- You generate code or published content at scale. The indemnity conditions matter more than the headline. Microsoft's required evaluation report on Azure OpenAI is work you must schedule, and Bedrock's uncapped cover only helps on Amazon's models.
- You need frontier models under ZDR. Covered Models on Anthropic's API require 30-day retention unless Anthropic authorizes otherwise, and Bedrock retains traffic on its named frontier models by default, with ZDR only by request or, for Claude Fable, through a program that ends December 31, 2026. The frontier tier and ZDR point in different directions on two of these platforms.
- The feature acts, not just answers. Evaluation evidence on your deployment outranks every certificate. Incident-notification clauses written for breaches may not fire for model misbehaviour, either.
What to Do Before the Next AI Renewal
This Week:
- Open the Microsoft 365 admin center under Copilot > Settings > AI providers operating as Microsoft subprocessors and record who has Anthropic and OpenAI enabled, and since when. If nobody decided that, it is your first finding.
- Send the six questions, verbatim, to every AI vendor due for renewal in the next two quarters, and require written answers with links. A vendor that answers with a SOC 2 bridge letter has told you something.
- Subscribe to the subprocessor notifications for OpenAI, AWS and Microsoft's Service Trust Portal. Put a monthly reminder on Anthropic's trust center, whose addendum does not name a notice period.
This Month:
- For each SaaS vendor with an AI feature, pull its SOC 2 system description and list the carved-out subservice organizations. Where a model provider is carved out, request that provider's own report and confirm the AI feature is in its scope.
- Run an indirect prompt-injection test against the assistant that reads mail — a crafted inbound message, EchoLeak-style — and keep the report. It is security evidence and, on Azure OpenAI, a condition of your copyright cover.
- Write down your retention ceiling per vendor as the maximum, not the default: two years for flagged content on Claude, 30 days for Fable-class traffic on Bedrock, 180 days for Google's deletion commitment. That is the number your records schedule has to accept.
Before Renewal:
- Ask Anthropic for a fixed subprocessor notice period, a termination right on an unresolved objection, and a written maximum for flagged-content retention. Ask Microsoft to publish the duration it holds abuse-monitoring samples for human review.
- Add base model, model publisher and licence to your standard AI clause, with written notice when any of them changes.
- Map each approved use case to the model route it actually takes, and confirm the indemnity and the certificate both follow that route.
The Bottom Line
Cloud security reviews went through this a decade ago. Buyers learned that "the provider is certified" and "my workload is secure" are different sentences, and the shared-responsibility model became the vocabulary for the gap. AI reviews are at the same point. This time the gap is not the configuration you own — it is the route your prompt takes after it leaves your tenant, and the terms that govern each hop.
The certificates are real, and the vendors here are mostly doing the work. But if your approved assistant is Copilot, it gained two outside model providers as subprocessors this year, and nothing in a clean SOC 2 opinion would have told you. The six answers would have.
A certificate tells you the vendor has a process. The six answers tell you what that process will do with your data.
Continue Reading
- Fable 5 Needs Retention On. ZDR Was Never Zero.
- Sony Sued Anthropic. Re-Prompting Voids Your Indemnity.
- CoCounsel's New Model Runs on Qwen. Go Read the Card.
- OpenAI Cuts Cursor Off Nov 12. BYOK Voids Your ZDR.
- 7 OpenAI Alternatives. Only 3 Clear a Sovereignty Rule.
- Stripe Bought OpenRouter. A Toggle Is Not a Contract.
