Agent Audit Logging: Platform Logs Miss the Decision You Must Prove

Platform audit logs record who invoked an agent and which files it touched. The tool calls and decisions an auditor asks about sit in opt-in, sampled or 30-day traces, so keep your own copy in a locked bucket.

By Rajesh Beri·October 3, 2026·18 min read
Share:
A steel archive cabinet in a server room with one drawer pulled open, filled with neatly labelled log tapes and a padlock hanging from the drawer handle, a laptop on top showing rows of timestamped records.

Illustration generated using AI

If an auditor asks you in March what your claims agent did on a Tuesday last October, the platform audit log will tell you who invoked it and when. It will almost never tell you which tool the agent called, with what arguments, under which policy, or what it saw before it acted. Every major platform splits the record in two: a durable, vendor-run audit log of who touched what, and a short-lived or opt-in trace of what the agent actually did. You have to join them, and you have to keep the second half yourself.

The recommendation: use each platform's native audit log for identity and admin events, emit OpenTelemetry agent spans for the decision trail, and land both in a write-once bucket you control, set to at least six months and more realistically your records policy. The native logs are free or nearly free and nobody on your team can quietly edit them. The traces are where the evidence lives, and every vendor-hosted copy of them can expire, be deleted by a user, or be turned off by default. Prices below were checked on each vendor's live page on October 3, 2026.

Platform What the audit record holds Prompt and tool content? Default retention Who can delete Price (checked Oct 3, 2026)
Microsoft Purview Audit (Copilot, Copilot Studio, Foundry) User, agent ID and version, files accessed, plugins, model provider Message IDs only, no text 180 days (Standard); 1 year with E5; up to 10 years with add-on Admins with retention-policy rights Included for Microsoft apps; $15 per 1M records for non-Microsoft AI apps
AWS: CloudTrail + Bedrock invocation logging + AgentCore Observability Caller ARN, model ID; full request and response if invocation logging is on Only with invocation logging or spans, both opt-in CloudTrail history 90 days; S3 and log groups are yours to set You; nobody, under S3 Object Lock compliance mode Data events $0.10 per 100k; S3 $0.023/GB-month
Google Cloud Audit Logs + request-response logging Admin Activity always on; Data Access off by default Only via request-response logging to BigQuery, sampled _Required 400 days; _Default 30 days, configurable to 3,650 Nobody, once a log bucket is locked 50 GiB per project per month free; retention past default billed
Anthropic Compliance API (Claude Enterprise) Activity metadata; chats and Claude Code session transcripts Not for API-key traffic; no thinking blocks or tool definitions Activity Feed 6 years; content per your policy Users (their chats), API callers (hard delete) Part of Claude Enterprise; pricing through sales
OpenAI Compliance Logs Platform (ChatGPT Enterprise/Edu) Audit, authentication and conversation logs as JSONL Yes, for ChatGPT workspace activity 30 days at OpenAI OpenAI expires it Part of ChatGPT Enterprise; pricing through sales
Langfuse (self-hosted or cloud) Full agent traces: prompts, tool calls, outputs, scores Yes, whatever you instrument Hobby 30 days, Core 90 days, Pro 3 years Project owners and admins Pro $199/month + $8 to $6 per 100k units; self-host MIT

The loser for this specific job is relying on Microsoft Purview alone for agent reconstruction, for reasons covered below. The pick for the trace half is OpenTelemetry into a store you own, with Langfuse or AgentCore Observability as the place your engineers read it.

What Does an Auditor Actually Ask For?

An agent audit record is the set of events that lets a third party reconstruct, after the fact, who asked an agent to act, what the agent decided, what it touched, and which control allowed it. That is four questions, and most logs answer only the first.

The minimum event set that survives a real inquiry:

  1. The actor pair: the human principal the agent acted for, and the agent's own identity (workload credential, agent ID, version). Purview's AgentId and AgentVersion fields and Anthropic's actor.user_id cover this half well.
  2. The input: the user prompt, the system prompt version (a hash is enough if the text is versioned elsewhere), and the retrieved context or a pointer to it.
  3. The model: provider, model ID and version per call. Agents switch models mid-run, and an auditor asking "was this the model you validated?" needs the answer per step.
  4. Every tool call: tool name, arguments, target resource, result status. This is the event most often missing.
  5. The authorization decision: which policy allowed or blocked the call, and any human approval with the approver's identity and timestamp.
  6. Guardrail interventions: what was filtered or blocked, and on which side (input or output).
  7. The output that left the system, and where it went.

OpenTelemetry's GenAI conventions already name the spans for this: invoke_agent for a full run and execute_tool for each tool call. Two caveats from the spec itself: the agent spans carry a "Development" stability status, and message content attributes are opt-in because they may contain personal data. Pin the convention version you emit, and decide content capture deliberately rather than inheriting a library default.

What Today's Platforms Actually Emit

The pattern across vendors holds: the always-on log records identity and resource access, and the content that explains a decision is opt-in, sampled, short-lived or absent.

Microsoft Purview Audit

Purview logs Copilot, Copilot Studio and Foundry agent interactions as part of Audit (Standard) with no extra setup, per Microsoft's audit documentation. The record is good on the "who and what was touched" half: AccessedResources lists the files, emails and sites the agent read, with sensitivity label IDs, an action value (read, create, modify) and a flag for detected cross-prompt injection. It also carries the plugin list and, for Copilot Chat, the model provider and name.

The Messages field holds message IDs and an isPrompt flag. The prompt and response text are not in the audit record. Retention is tied to per-user licensing: Standard is 180 days, one year needs an E5 or Purview Suite license on the user who generated the record, and three to ten years needs the 10-Year Audit Log Retention add-on on top. Non-Microsoft AI apps audited through Purview are billed pay-as-you-go at $15 per 1 million records ingested and kept 180 days.

The episode that should shape how much you trust any single vendor log: in July 2025 a researcher found that asking Microsoft 365 Copilot to summarize a file without returning a reference link produced no audit entry. Microsoft fixed it, classified it "important", and told customers nothing, The Register reported. The researcher's summary: "Your audit log is wrong, and Microsoft doesn't plan on telling you that."

AWS: CloudTrail, Bedrock Invocation Logging, AgentCore

AWS gives you the most raw material and turns most of it off. CloudTrail logs InvokeModel and Converse as management events, on by default, recording the caller's ARN and the model ID. The documented example has responseElements: null: no prompt, no output. InvokeAgent, knowledge base retrieval and ApplyGuardrail are data events, which CloudTrail does not log until you add advanced event selectors. AWS's own note on guardrail events is worth reading before you rely on them: when several guardrails evaluate one call, "the CloudTrail event doesn't identify which guardrail produced each assessment."

The content lives in Bedrock model invocation logging, which captures the full request and response body (inline up to 100 KB, larger bodies as separate S3 objects) plus the caller ARN and optional request metadata tags. It is "disabled by default", and it covers only the bedrock-runtime endpoint. For agents hosted on AgentCore Runtime, AgentCore Observability instruments the agent with OpenTelemetry automatically, but spans are only searchable after you enable CloudWatch Transaction Search, whose console setup offers to index 1% of traces at no cost. A sampled trace is a debugging aid. It fails an audit the first time the run you need was in the other 99%, so for audit purposes route the raw spans to your own log group and archive them in full.

The Amazon Bedrock pieces are the most complete set of primitives here. They are also the most assembly: three services, two opt-ins, one per-account switch.

Google Cloud Audit Logs and Request-Response Logging

Google's Admin Activity audit logs "are always written; you can't configure, exclude, or disable them." Data Access logs, which record reads and writes of user data, are "disabled by default because they can generate large volumes of data." Neither carries prompt text. For content, Vertex AI request-response logging writes prompts and responses to a BigQuery table you name, with a samplingRate between 0 and 1. Set it to 1 for anything you may have to defend.

Retention is clear: the _Required bucket keeps audit logs 400 days and is not configurable; _Default keeps 30 days in a project, configurable from 1 to 3,650 days. Google's pricing page includes 50 GiB of ingestion per project per month and bills retention beyond the default period. For Google Vertex AI agents, the strongest feature is the locked log bucket, covered under tamper evidence below.

Anthropic Compliance API

For Claude Enterprise, the Compliance API's Activity Feed has the longest default retention of any vendor record in this comparison: activities are queryable within a minute and retained for 6 years. Session transcripts for Claude Code and Cowork include user prompts, assistant responses and tool activity.

Anthropic's own documentation lists the gaps, which is more than most vendors do. The Compliance API does not include prompt text or responses "from Claude Console, or from Claude API workloads authenticated with an API key." It omits thinking blocks, tool definitions and MCP server configuration from transcripts, and Claude Code run through Bedrock, Google Cloud or Foundry. Recording "is not retroactive": nothing before you enable it is backfilled. Chats a user deletes stay listed, but their content is gone, and a Compliance API hard delete is "immediate and permanent." A practitioner review from Monad puts it plainly: the feed shows that something happened, and does not show "intent, legitimacy, or impact." If your agents call the Claude API with a key, which is how most custom agents run, this API holds none of their content and your own traces are the only record.

OpenAI Compliance Logs Platform

For ChatGPT Enterprise and Edu, OpenAI's Compliance Platform exports audit, authentication and conversation logs as immutable, append-only JSONL files. The catch is the window. Integrators such as Elastic build around the 30-day retention window: anything you have not downloaded in 30 days is gone. On the API side, OpenAI's data controls page says abuse monitoring logs are kept up to 30 days, and Responses API state is stored 30 days by default. Neither is an audit record you control. Treat the platform as a feed to drain on a schedule.

Langfuse

Langfuse is the reference choice for the trace half when you want one store across clouds and model vendors. It captures everything you instrument, and its plans price that by units. Pro costs $199 a month with 100,000 units, then $8 per 100,000 up to 1M and $7 per 100,000 up to 10M, with three years of data access. The self-hosted edition is free under the MIT licence, except the ee/ directories, and configurable data retention needs the Enterprise Edition when self-hosted.

That retention feature is also why Langfuse is the wrong final resting place for audit evidence. Project owners and admins set a retention window, and "on a nightly basis, Langfuse selects traces, observations, scores, and media assets that are older than the configured retention period and deletes them." An admin who can shorten that window can remove evidence. Run Langfuse for engineers and export to WORM storage for auditors. (We compared the observability vendors on billing in an earlier teardown.)


How Much Does Keeping It Cost at Agent Volume?

At agent volume the archive is cheap and the vendor-hosted, queryable retention is what costs money. To compare like for like, take one workload: 50 production agents, 200,000 runs a month, 25 spans per run (model calls plus tool calls), about 4 KB per span with content captured. That is 5 million spans and roughly 20 GB a month before compression.

  • Raw archive on S3: 20 GB a month at $0.023 per GB-month for S3 Standard adds about $0.46 to the bill each month. After seven years you hold about 1.7 TB, roughly $39 a month at Standard rates, and less if you tier it down.
  • AWS ingestion: logging 200,000 InvokeAgent calls as CloudTrail data events costs about $0.20 at $0.10 per 100,000 events. Sending 20 GB of spans through CloudWatch Logs at its $0.50 per GB standard ingestion rate is about $10.
  • Purview: Copilot Studio and Foundry agents are inside Audit Standard at no extra charge. The cost is the license that sets retention: one year needs E5 or Purview Suite on each user generating records, and longer needs the 10-year add-on, a separate license purchase.
  • Langfuse Cloud Pro: 200,000 traces plus 5 million observations is about 5.2 million units. That is $199, plus $72 for units up to 1M, plus about $294 for the next 4.2M, or roughly $565 a month for three years of queryable access.

At this workload, archive storage is a rounding error and the hosted trace tool is where the money goes, so keep the hosted tool's retention short (90 days covers most debugging) and keep the long tail in object storage, where a seven-year archive costs less than one month of the hosted tool.

Who Can Delete Your Agent Logs?

Tamper evidence means two things: no one on your team can alter or delete a record before its retention ends, and you can show that a given export is complete. Most stacks fail the first test because the same admin who runs the agent platform can change its logging.

The controls that pass:

  • S3 Object Lock in compliance mode: a locked object version "can't be overwritten or deleted by any user, including the root user", and the only way to remove it early is to close the AWS account. Governance mode is weaker: anyone with s3:BypassGovernanceRetention can delete, and the S3 console sends the bypass header by default. Use governance mode to test, then switch.
  • Locked Google log buckets: locking is irreversible, and the bucket cannot be deleted until every entry has served its retention.
  • Provenance on every export: Anthropic's guidance for its own feed generalizes: store "source endpoint, query parameters, run timestamp, and a content hash of each record," and log the starting cursor, the terminal ID and the record count per run. The feed is at-least-once, so deduplicate on the activity ID.

The controls that do not pass: a Langfuse project with a retention window any admin can shorten, a CloudWatch log group whose retention setting the platform team owns, a vendor feed that expires in 30 days, and a Purview retention policy that an admin with the Organization Configuration role can delete. Keep them as working copies and put the record somewhere else.

What the EU AI Act and SOC 2 Expect

The EU AI Act sets the only hard number here, and it applies to high-risk systems. Article 12 requires that high-risk AI systems "technically allow for the automatic recording of events (logs) over the lifetime of the system." That obligation sits with the provider. As a deployer, Article 26(6) requires you to keep the logs under your control "for a period appropriate to the intended purpose" of "at least six months." The Digital Omnibus moved the start date: stand-alone (Annex III) high-risk obligations now apply from 2 December 2027, and embedded ones from 2 August 2028. If your agent screens job applicants or scores credit, that is the deadline to design against.

Map it against the table and the gaps show. OpenAI's 30-day window and Google's 30-day _Default bucket both fall short of six months unless you export. Purview's 180 days for Standard just clears it. Anthropic's six years clears it for metadata, and for API-key agents there is no content record at all.

SOC 2 is less prescriptive. CC7.2 requires ongoing monitoring of system components to detect anomalies, and names no retention period. In practice your auditor samples events from across the audit period and asks you to produce the logs, so retention has to cover the period plus the time to fieldwork. For a 12-month period with a month of fieldwork, that means 13 months. The test your agent logs will face is whether the sampled tool call can be traced to a human principal and an authorization decision.

Who Should Not Pick Each Option

  • Purview as your agent record: skip it if your agents act outside Microsoft 365, if you need prompt text in the audit record itself, or if your users are not all on E5. It is a good resource-access log for Copilot and nothing more.
  • AWS primitives: skip them if no one on your team will own three opt-ins across accounts and Regions. Invocation logging misses calls to the bedrock-mantle endpoint, so check which endpoint your SDK uses.
  • Google Cloud logging: skip it if you cannot afford a 100% sampling rate on request-response logging for regulated agents, because a sampled log cannot produce a specific run on request.
  • Anthropic Compliance API: do not count on it for custom agents calling the Claude API with keys. It covers Claude Enterprise apps and Claude Code sessions, not your API workloads.
  • OpenAI Compliance Logs Platform: do not use it as an archive. If nobody owns the job that drains it every day, a long enough outage in that job leaves a permanent hole in your record.
  • Langfuse: do not make it the system of record, and skip self-hosting it if you need retention controls without buying the Enterprise Edition.

Which Option Loses, and Why?

For reconstructing what an agent did, Purview alone loses. It has the best structured resource-access record of the six, with sensitivity labels and injection flags that the others lack. But the record holds message IDs instead of text, its retention depends on what license each user holds, and its best-documented failure was a silent gap Microsoft did not disclose. If your agents run on Copilot Studio, keep Purview and add traces. Do not let a clean Purview export convince anyone the audit question is answered.

Decision Criteria That Predict Regret

Buyers' regret in this category usually traces to one of five questions, answered late:

  1. Is the tool call in the durable record? If the answer depends on an opt-in setting, put that setting in infrastructure-as-code with a drift alert.
  2. Who can shorten retention? If it is the same team that runs the agents, you do not have tamper evidence.
  3. Is any layer sampled? Transaction Search at 1% and request-response logging below 1.0 are both sampling.
  4. Does the record survive a user deleting their chat? For Claude Enterprise chats and remote sessions, it does not, unless you exported first.
  5. Can you join the halves? A run ID that appears in the platform audit log, your trace and your authorization log is the difference between a two-hour answer and a two-week one. Use the OTel trace ID and pass it as Bedrock request metadata or a span attribute.

What changes the answer: if all your agents run inside one ecosystem (Copilot Studio only, or Claude Enterprise only), the native record plus a daily export to WORM storage may be enough. The moment one workflow crosses two vendors, you need OpenTelemetry as the common format, because no vendor's audit log will show the other vendor's half.

What to Do Next

This Week:

  1. Pick one production agent and try to reconstruct a run from 30 days ago using only what is stored today. Write down which of the seven events you could not find.
  2. Check three settings: Bedrock model invocation logging, CloudTrail data event selectors for AWS::Bedrock::AgentAlias, and Vertex request-response samplingRate. All three default to off or partial.

This Month:

  1. Create an S3 bucket with Object Lock in compliance mode (or a locked Google log bucket) with retention of at least 13 months, and route agent spans, invocation logs and vendor compliance exports to it.
  2. If you are on ChatGPT Enterprise, schedule a daily export from the Compliance Logs Platform and alert when it fails. If you are on Claude Enterprise, enable the Compliance API now, since nothing before enablement is recorded.

Before Q1 Close:

  1. Agree a single run ID across agent framework, gateway and identity logs, and show your auditor one sampled tool call traced end to end.
  2. For any agent that could fall under Annex III, write down its log retention period and owner against Article 26(6) before the December 2027 date.

The Bottom Line

This is the same split that happened with cloud infrastructure a decade ago. CloudTrail told you who called the API; it did not tell you what the application did with the result, so teams built application logging and shipped it somewhere the platform admin could not touch. Agent platforms have repeated the pattern, with the decision trail in the opt-in, short-lived layer. Pay for the hosted trace tool that your engineers will use, keep its retention short, and write everything to a locked bucket that costs about $39 a month after seven years at 200,000 runs a month.

Turn on Bedrock invocation logging and set Vertex sampling to 1.0 before the next agent ships.

Continue Reading

Share:

Frequently Asked Questions

What should an AI agent audit log contain?

At minimum: the human principal and the agent's own identity and version, the prompt and system prompt version, the model per call, every tool call with arguments and target, the authorization decision or human approval, guardrail interventions, and the output. Most platform audit logs cover only identity and resource access.

How long does the EU AI Act require agent logs to be kept?

Article 26(6) requires deployers of high-risk AI systems to keep logs under their control for at least six months. After the Digital Omnibus, stand-alone high-risk obligations apply from 2 December 2027 and embedded ones from 2 August 2028.

Does Microsoft Purview log Copilot Studio agent prompts?

Purview's CopilotInteraction record includes the agent ID and version, the files and sites accessed, plugins and message IDs, but not the prompt or response text. Standard retention is 180 days; one year needs E5 and up to ten years needs an add-on license.

How do I make agent logs tamper-evident?

Write them to storage nobody on your team can alter, such as S3 Object Lock in compliance mode, where even the root user cannot delete a locked object, or a locked Google Cloud log bucket. Record a content hash and run metadata for every export.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →