AgenticModels

How we built Claude Code auto mode: a safer way to skip permissions

by Anthropic

IntermediateGuideFree~25 min read

The design, threat model and measured error rates of the classifier that decides which Claude Code actions run without asking you.

Start LearningAdded Oct 4, 2026 · Updated Oct 4, 2026

Overview

Published on March 25, 2026 and written by John Hughes with eleven credited contributors, this post explains the permission system behind Claude Code auto mode. Its sections cover how it works, the threat model, how permission decisions work, the classifier decision criteria, results, design decisions and next steps. The threat model names four failure types: overeager behaviour, honest mistakes, prompt injection and a misaligned model. Actions fall into three tiers: a built-in allowlist of read-only tools, in-project file edits that skip the classifier because version control can review them, and everything else, which goes to a transcript classifier running on Claude Sonnet 4.6. That classifier sees only user messages and tool call payloads; the agent's own messages and tool outputs are stripped so it cannot be talked into approving something. Stage one is a fast single-token filter and stage two adds chain-of-thought reasoning only when stage one flags an action. Evaluated on 10,000 real internal tool calls, 52 curated overeager actions and 1,000 synthetic exfiltration attempts, the full pipeline shows a 0.4% false positive rate, a 17% false negative rate on overeager actions and 5.7% on exfiltration. The post says auto mode is not a replacement for careful human review on high-stakes infrastructure, and explains the deny-and-continue fallback, which escalates after 3 consecutive or 20 total denials.

At a Glance

Topic
Agentic
Level
Intermediate
Format
Guide
Cost
Free
Duration
~25 min read
Provider
Anthropic
Hands-on
No
Certificate
None

What You’ll Learn

  • ✓Classify agent risks into overeager behaviour, honest mistakes, prompt injection and misaligned model goals
  • ✓Split tool permissions into a read-only allowlist, in-project edits and classifier-reviewed actions
  • ✓Design a classifier that sees user messages and tool calls but not the agent's own reasoning
  • ✓Use a cheap single-token first stage and a reasoning second stage to trade recall against false positives
  • ✓Write block rules and allow exceptions around a declared trust boundary of domains, buckets and services
  • ✓Handle a blocked action with deny-and-continue and escalate to a human after repeated denials
  • ✓Check delegated and returned work in multi-agent handoffs for unauthorised tasks and injected instructions

Highlights

  • •Publishes its own error rates per stage, including the 17% false negative rate on real overeager actions and why prompt engineering did not fix it
  • •Lists real blocked cases such as scope escalation on branch cleanup, credential hunting after an auth failure and a deploy retried with a skip-verification flag
  • •An independent arXiv stress test (2604.04978) measured an 81.0% end-to-end false negative rate on ambiguous authorization prompts and found 36.8% of state-changing actions went through unreviewed in-project edits
  • •An Embrace The Red report from August 2026 bypassed auto mode in 60-80% of small-sample runs; Anthropic replied that the classifier is best-effort and OS isolation is the security boundary
  • •The command claude auto-mode defaults prints the full default rule set the post describes

Who It’s For

Best For

  • ✓Engineers building permission or approval layers for tool-using agents
  • ✓Security teams deciding whether to allow coding agents to run unattended
  • ✓Claude Code users who currently run with --dangerously-skip-permissions
  • ✓Developers writing guardrails for prompt injection in agent tool outputs

Prerequisites

  • •Familiarity with coding agents and tool calling
  • •Basic understanding of prompt injection and classifier precision and recall

FAQ

What is How we built Claude Code auto mode: a safer way to skip permissions?

An Anthropic engineering post by John Hughes explaining how Claude Code auto mode replaces permission prompts with a prompt injection probe and a two-stage transcript classifier. It is for engineers building agent permission systems or deciding whether to run coding agents unattended, and it publishes false positive and false negative rates.

Is How we built Claude Code auto mode: a safer way to skip permissions free?

How we built Claude Code auto mode: a safer way to skip permissions is free to access.

What level is How we built Claude Code auto mode: a safer way to skip permissions for?

How we built Claude Code auto mode: a safer way to skip permissions is aimed at a intermediate audience. Recommended background: Familiarity with coding agents and tool calling, Basic understanding of prompt injection and classifier precision and recall.

How long does How we built Claude Code auto mode: a safer way to skip permissions take?

Expect roughly ~25 min read. Most learners work through it at their own pace.

What will I learn from How we built Claude Code auto mode: a safer way to skip permissions?

You'll learn: Classify agent risks into overeager behaviour, honest mistakes, prompt injection and misaligned model goals; Split tool permissions into a read-only allowlist, in-project edits and classifier-reviewed actions; Design a classifier that sees user messages and tool calls but not the agent's own reasoning; Use a cheap single-token first stage and a reasoning second stage to trade recall against false positives; Write block rules and allow exceptions around a declared trust boundary of domains, buckets and services; Handle a blocked action with deny-and-continue and escalate to a human after repeated denials; Check delegated and returned work in multi-agent handoffs for unauthorised tasks and injected instructions.

Topics

claude-codeagent-securitypermissionsprompt-injectionguardrailsclassifier

Sources

This page was written from 4 sources, 3 on domains other than anthropic.com.

  1. 1.anthropic.com — claude code auto modevendor
  2. 2.arxiv.org — 2604.04978
  3. 3.embracethered.com — breaking claude code opus 5 and automode
  4. 4.hn.algolia.com — search