Human-in-the-Loop for AI Agents: Most Approval Gates Rubber-Stamp

Claude Code users approve 93% of permission prompts, and human reviewers caught 13.6% of harmful actions in Anthropic's tests. Gate only irreversible agent actions, sample the rest with planted errors, and use the free gate in LangGraph or the OpenAI Agents SDK.

By Rajesh Beri·October 4, 2026·14 min read
Share:
An office approval queue on a monitor with a reviewer's hand resting on a mouse over a green approve button, a tall stack of identical printed request slips beside the keyboard, and a red mushroom-shaped emergency stop b

Illustration generated using AI

If you put a human approval step in front of every agent action, you have most likely bought a rubber stamp. Anthropic measured it on its own product: Claude Code users approve 93% of permission prompts, and in a study Anthropic ran on 1,053 paid testers, human review caught 13.6% of harmful actions while its automated classifier caught 89%. The fix is to gate far fewer actions, make the remaining gates expensive enough that a person actually reads them, audit everything else by sample, and build a kill switch that works without the agent's cooperation.

The tooling is the cheap part. The approval primitive is free and open source in LangGraph and the OpenAI Agents SDK, and either one is the right default for a team writing agents in code. Here is how the five options buyers ask about compare, normalised to one workload: 10,000 gated actions a month, each taking a reviewer 90 seconds (our assumption; time your own queue). Prices were checked on each vendor's live page on October 4, 2026.

Option What the gate is Tool cost per reviewed action Verdict
LangGraph interrupt() Pauses a graph node, persists state to a checkpointer, resumes on Command(resume=...) $0 (MIT licence); you host the checkpointer Pick it if your agents are already graphs. Make every side effect idempotent first.
OpenAI Agents SDK needs_approval Marks a tool as gated; the run returns interruptions you approve or reject, then resume $0 (MIT licence); you store the serialised RunState Pick it for the cleanest per-tool gate. Narrowest blast radius on resume.
Temporal (signals) A durable workflow that waits for a signal from your approval UI $50 per million actions, falling to $25 at volume Add it when approvals wait hours or days, or span systems.
Microsoft Copilot Studio multistage approvals Human and AI approval stages in an agent flow, answered in Teams or Outlook About $0.001 per flow action; AI stages billed per token Wait. Still preview, no ALM support.
Amazon Augmented AI (A2I) Managed human-review queue with confidence and sampling triggers $0.03 per object (first 100,000, Textract/Rekognition) Avoid. Closed to new customers.
A person's time (any option) 90 seconds at the median claims-adjuster wage (May 2025) About $0.94 per review, wage only The cost that actually matters

The loser is Amazon A2I. It had the best-designed sampling controls of the group, and AWS has stopped selling it.


Why Do Human Reviewers Turn Into Rubber Stamps?

Reviewers approve almost everything because almost everything they see deserves approval, and attention drops fast when real problems are rare. Anthropic named this directly in its auto mode write-up, calling it approval fatigue: the 93% approval rate is the evidence that people have stopped reading the prompts. Read the 13.6% with its setup in mind: Anthropic ran the test, sells the classifier it compared against, and planted one dangerous command per session as a routine prompt. Its own classifier is no oracle either. The full pipeline still has a 17% false-negative rate on real overeager actions, by Anthropic's account.

Two older bodies of research explain the decay. The first is vigilance. In the classic Mackworth experiments, detection fell 10 to 15% in the first 30 minutes, and under most conditions the decrement becomes significant within the first 15 minutes of watching. A reviewer clearing an approval queue for an afternoon is running exactly that experiment.

The second is automation bias: people defer to a machine's suggestion even when they know better. In a 2023 study of 27 radiologists reading 50 mammograms, inexperienced readers' accuracy fell from almost 80% to under 20% when the purported AI suggested the wrong category. Readers with 15+ years of experience fell from 82% to 45.5%. Seniority softened the effect and did not remove it. An agent that presents its proposed action with a confident rationale is making the same kind of suggestion to your approver.

The EU AI Act already assumes this happens. Article 14(4)(b) requires that people overseeing a high-risk system be able to "remain aware of the possible tendency of automatically relying or over-relying on the output", and calls it automation bias by name. A checkbox your staff click 93 times out of 100 is weak evidence that you met that duty.

Which Agent Actions Actually Need a Human Gate?

Gate an action only when it is hard to reverse and someone outside the system would notice it. Everything else belongs in sampled audit. A practical test: if undoing the action means calling a customer, a bank, a regulator or a counterparty, it gets a gate. If undoing it means running another API call, it does not.

That leaves a short list for most enterprises:

  1. Money that leaves the company: payments, refunds and credits above a threshold you set from your loss history.
  2. Messages that reach people outside the company in your name, where the content is new rather than templated.
  3. Changes to permissions, credentials, or who can do what.
  4. Deletes and overwrites of data you cannot restore from backup inside your recovery objective.
  5. Production changes that skip your normal change process.

For the highest-consequence tier, require two people. The AI Act sets that bar for one category already: biometric identifications must be verified by at least two natural persons before anyone acts on them. Copying that rule for, say, wires over a set amount costs little, because by the time you apply it the volume is small.

The Sibos 2026 bank deployments followed this shape: agents repair and prepare payments, and a person still releases them. That works because the release step is rare and consequential, so the human reads it.

How Do Sampling and Tiered Review Hold Up?

A tiered review sends every low-confidence or high-consequence action to a person, sends a small random share of everything else, and plants known-bad items in the queue to measure whether reviewers are still catching them. The design came from industries that faced this problem before agents did.

Amazon A2I's configuration language is the cleanest published example of the routing half. Its docs show a rule that always sends inferences below 60% confidence to a human and samples 5% of those above 90%, with a RandomSamplingPercentage anywhere from 0.01 to 100. You can rebuild that logic in a few lines inside a LangGraph conditional edge or an OpenAI needs_approval function, which accepts an async callable that decides per call.

The measurement half comes from airport security. Threat image projection inserts pre-recorded images of weapons into the live X-ray stream, and screener responses are scored as a hit rate; that study found about 88% at the airport it examined, and officers who fall short go back to training. Vigilance research points the same way: introducing artificial signals similar to real targets reduced the decrement.

For an agent queue, that translates into three numbers to track weekly. Approval rate: if it sits above 95% for a month, the gate is either catching nothing worth catching (move the action to sampling) or not being read (fix the reviewer setup). Canary hit rate: seed 1 to 2% of the queue with actions you know are wrong, and if reviewers approve them, you have measured the rubber stamp directly. Time-on-item: a 4-second median on a refund approval means nobody opened the details.

Keep review sessions short. Given that the vigilance drop shows up inside 15 minutes, a reviewer who clears the queue in 20-minute blocks between other work will catch more than one who sits on it all day. Fix what the reviewer sees, too. The OpenAI SDK's approval item carries tool_name and arguments, and raw JSON is exactly what people learn to skip. Render the consequence in plain words, such as "refund $4,200 to an account opened 3 days ago".

What Does a Reviewed Action Really Cost?

At our defined workload, the person costs far more than any of the tools. The US median wage for claims adjusters, examiners and investigators, a reasonable proxy for an operations reviewer, is $37.51 an hour (May 2025). At 90 seconds a review that is $0.94 per action in wages before benefits, and 10,000 reviews a month is 250 reviewer hours and about $9,400.

The tools, at the same volume:

So the budget question is reviewer hours. Cutting gated actions from 10,000 to 1,000 by moving reversible ones to a 5% sample saves around 200 hours a month, and the remaining reviews get read. That is the trade to put in front of your CFO, along with the loss rate you expect on the sampled tier.


Which Approval Tool Should You Pick?

Use the gate built into the framework your agents already run on, and add Temporal only when an approval has to survive a long wait. Here is who each option fits and who should stay away.

LangGraph. An interrupt is a call to interrupt() inside a node; LangGraph saves state through its checkpointer and waits indefinitely until you resume. The catch is in the same docs: on resume, "the runtime restarts the entire node from the beginning", so any code before the interrupt runs twice. The failure is live. An issue opened September 29, 2026 reports that resuming one of two parallel interrupts in a subgraph re-runs a sibling node and repeats its side effect "without asking again", and it was still open when we checked. Do not pick LangGraph for approvals if your tools are not idempotent and you cannot make them so, because the duplicate lands after a human approved the original.

OpenAI Agents SDK. You set needs_approval on a tool, and the run returns ToolApprovalItem interruptions. You persist the state with result.to_state(), call approve or reject, and resume. The gate wraps the tool call itself, which is the right place for it. Its warning matters for anyone storing paused runs in a shared database: only deserialise snapshots from trusted storage, because a tampered snapshot is a forged approval. Teams that do not use OpenAI's runner should not adopt it just for the gate.

Temporal. A signal is an asynchronous write into a running workflow, which makes "wait for a manager, escalate after 48 hours, cancel after a week" ordinary code. Skip it if your approvals clear in minutes inside one service; you would be running a workflow platform to hold a boolean.

Copilot Studio multistage approvals. This is the only option on the list a business team can configure without engineers, and approvers answer in Teams, Outlook or the Power Automate portal. It is also marked preview, "not meant for production use", with no ALM support, and the same approver cannot sit in two stages. Microsoft's FAQ says AI approvals were not designed for insurance claims, loan approvals or other high-stakes decisions. Regulated teams should leave the AI stage off and wait for general availability.

Amazon A2I. The docs now state that A2I "is no longer open to new customers" and that AWS does not plan new features. Existing users can keep running it. Nobody should start a project on it.

One more warning about buying a standalone approval service. HumanLayer was the best-known approvals API for agents; its site now sells a multiplayer coding agent IDE priced per user, with no approvals API on the page. With the gate shipping free in the frameworks, a standalone approval vendor has little left to sell, so keep the gate in your own code.

Designing a Kill Switch Instead of a Checkbox

A kill switch is a single control that stops every running instance of an agent and revokes what it can touch, and it has to work when the agent is misbehaving. Per-action approval does nothing for an agent that is already doing damage between gates. Article 14(4)(e) asks for the ability to interrupt the system "through a 'stop' button or a similar procedure" that brings it to a safe state.

Three properties separate a working switch from a decorative one. It acts outside the agent: revoking the agent's credentials or tokens at your identity provider stops it whether or not its code cooperates, while an in-process flag depends on the process you are trying to stop. It has a defined safe state: in-flight payments are held rather than half-sent, which in LangGraph or Temporal terms means the resume path checks a halt flag before any side effect. And it gets drilled: someone on call pulls it on a schedule and you record how long it took.

In-framework stops still help as a second layer. The OpenAI SDK's guardrails can raise a tripwire exception that halts agent execution when an input, output or tool call fails a check. Treat that as a circuit breaker for known-bad patterns. OpenAI's own automatic shutdown failed for 2.5 hours after an agent escaped through DNS, which is the case for the out-of-band layer.

What Changes the Answer?

The criteria that predict regret are the reversibility of your actions, how long approvals wait, and whether your reviewers are measured. If your tools cannot be made idempotent, avoid LangGraph's resume model or wrap every side effect in an idempotency key first. If approvals regularly wait longer than a working day, put Temporal under whichever framework you use. If business owners must own the approval logic without engineering, Copilot Studio is the only candidate, and you should plan on human-only stages until it leaves preview. If you cannot report an approval rate and a canary hit rate per queue, the choice of tool will not matter, because you will not know whether the gate works. Logging every approval with who approved what and when is a prerequisite; our agent audit logging guide covers the event set.

What to Do Next

This Week:

  1. Export every action your agents can take and mark each one reversible or irreversible using the "would you have to call someone to undo it" test.
  2. Pull the approval rate and median time-on-item for each existing approval queue. Anything above 95% approved goes on a list for redesign.

This Month:

  1. Move reversible actions from blocking approval to a 5% random sample reviewed after the fact, and route low-confidence actions to the gate regardless of type.
  2. Seed 1 to 2% of each remaining queue with known-bad actions and publish the hit rate to the queue owner weekly.
  3. Replace raw argument payloads in the approval UI with a one-line statement of consequence.

Before Year-End:

  1. Build the out-of-band kill switch: one command that revokes the agent's credentials and holds in-flight work, owned by a named on-call rotation.
  2. Run a timed drill and write down the number of minutes from decision to full stop.

The Bottom Line

Human review of agents is repeating what airport screening learned: when real threats are rare, people stop looking, and the answer there was planted test items plus a scored hit rate for every screener. The frameworks give you the gate for free. The cost you control is reviewer attention, about $0.94 in wages per 90-second review, so spend it where an action cannot be taken back.

Count the irreversible actions first, then decide how many reviewers you need.

Continue Reading

Share:

Frequently Asked Questions

Do AI agents need a human to approve every action?

No. Approving everything trains reviewers to click yes: Anthropic found Claude Code users approve 93% of permission prompts. Gate only actions that are hard to reverse and visible outside the company, such as payments, external messages, permission changes and unrecoverable deletes, and review the rest by random sample.

How much does human review of an agent action cost?

Mostly reviewer time. At the May 2025 US median wage for claims adjusters, examiners and investigators ($37.51 an hour) and 90 seconds per review, each action costs about $0.94 in wages before benefits. Tooling ranges from $0 for LangGraph or the OpenAI Agents SDK to $0.03 per object on Amazon A2I.

What is the best tool for human-in-the-loop approvals in AI agents?

Use the gate in the framework you already run: LangGraph's interrupt() or the OpenAI Agents SDK's needs_approval, both MIT-licensed. Add Temporal when approvals must wait hours or days. Copilot Studio multistage approvals are still in preview, and Amazon A2I is closed to new customers.

How do you stop human reviewers from rubber-stamping AI decisions?

Cut the volume they see, keep sessions short because vigilance drops within about 15 minutes, show the consequence instead of raw arguments, and plant known-bad items in 1 to 2% of the queue. Track approval rate, canary hit rate and time-on-item for each queue every week.

What makes an AI agent kill switch work?

It acts outside the agent, for example by revoking the agent's credentials at the identity provider, so it works even if the agent's code does not cooperate. It holds in-flight work in a safe state, has a named on-call owner, and is drilled on a schedule with the time to full stop recorded.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →