Quantifying infrastructure noise in agentic coding evals
by Anthropic
Why a two-point lead on an agentic coding leaderboard can come from container limits rather than the model.
Overview
Published on February 5, 2026 by Gian Segato of Anthropic, with contributions from Nicholas Carlini, Jeremy Hadfield, Mike Merrill and Alex Shaw, this post reports a controlled experiment on agentic coding evals. The team ran Terminal-Bench 2.0 on a Google Kubernetes Engine cluster under six resource configurations, from strict enforcement (guaranteed allocation equal to the hard kill limit) to fully uncapped, while holding the Claude model, harness and task set constant. Infrastructure error rates fell from 5.8% at strict enforcement to 0.5% uncapped, and success rose by 6 percentage points (p < 0.01). The key finding is a two-regime split: up to about 3x headroom, extra resources mostly remove spurious out-of-memory kills, and the score change stays within noise (p = 0.40). Above 3x, resources start letting agents solve tasks they otherwise could not, for example by installing heavy dependencies, so the benchmark measures something different. A crossover experiment on 227 SWE-bench problems with 10 samples each found a smaller +1.54 point effect at 5x RAM. The post closes with recommendations: specify guaranteed allocation and kill threshold separately per task, calibrate the gap between them, report the multiplier, and treat leaderboard gaps under 3 points with skepticism until configurations match. It also flags time-of-day pass-rate drift from API latency. Terminal-Bench 4.0 states that it calibrated its task resources using this methodology.
At a Glance
- Topic
- Agentic
- Level
- Advanced
- Format
- Guide
- Cost
- Free
- Duration
- ~20 min read
- Provider
- Anthropic
- Hands-on
- No
- Certificate
- None
What You’ll Learn
- ✓Distinguish guaranteed resource allocation from the hard kill threshold in container runtimes used for agent sandboxes
- ✓Explain why pinning allocation equal to the limit causes spurious failures from transient memory spikes
- ✓Recognize the 3x headroom boundary where extra resources stop fixing reliability and start making tasks easier
- ✓Configure per-task resource floors and ceilings so eval scores reflect model capability rather than infrastructure
- ✓Read agentic coding leaderboards skeptically when gaps are under three points and configurations are undocumented
- ✓Account for confounders such as time-of-day API latency when comparing agent pass rates across runs
- ✓Design a controlled resource-sweep experiment holding model, harness and task set constant
Highlights
- •Controlled experiment with statistical tests (p-values reported) across six resource configurations rather than anecdote
- •Cross-checked on a second benchmark (SWE-bench, 227 problems x 10 samples) to show the effect generalises at a smaller magnitude
- •Terminal-Bench 4.0's maintainers cite it as the methodology they used to calibrate task resources
- •Gives a concrete, citable rule of thumb (be skeptical of gaps under 3 points) for readers of model leaderboards
- •Independent write-ups such as the ZenML LLMOps Database rate the method rigorous, while arguing that resource limits are part of the real task rather than pure noise
Who It’s For
Best For
- ✓Engineers building or operating agent eval harnesses on Kubernetes or container sandboxes
- ✓ML teams choosing between models based on Terminal-Bench or SWE-bench leaderboard scores
- ✓Benchmark maintainers deciding how to specify and document task resource limits
- ✓Researchers designing reproducible agentic evaluation setups
Prerequisites
- •Familiarity with agentic coding benchmarks such as SWE-bench or Terminal-Bench and how pass rates are computed
- •Working knowledge of containers and resource requests versus limits (Docker or Kubernetes)
- •Basic statistics: confidence intervals and p-values
FAQ
What is Quantifying infrastructure noise in agentic coding evals?
An Anthropic engineering study for anyone who runs or reads agentic coding benchmarks such as Terminal-Bench and SWE-bench. It measures how container CPU and memory enforcement alone moves agent scores, and gives concrete rules for configuring eval sandboxes and for deciding which leaderboard gaps are real enough to act on.
Is Quantifying infrastructure noise in agentic coding evals free?
Quantifying infrastructure noise in agentic coding evals is free to access.
What level is Quantifying infrastructure noise in agentic coding evals for?
Quantifying infrastructure noise in agentic coding evals is aimed at a advanced audience. Recommended background: Familiarity with agentic coding benchmarks such as SWE-bench or Terminal-Bench and how pass rates are computed, Working knowledge of containers and resource requests versus limits (Docker or Kubernetes), Basic statistics: confidence intervals and p-values.
How long does Quantifying infrastructure noise in agentic coding evals take?
Expect roughly ~20 min read. Most learners work through it at their own pace.
What will I learn from Quantifying infrastructure noise in agentic coding evals?
You'll learn: Distinguish guaranteed resource allocation from the hard kill threshold in container runtimes used for agent sandboxes; Explain why pinning allocation equal to the limit causes spurious failures from transient memory spikes; Recognize the 3x headroom boundary where extra resources stop fixing reliability and start making tasks easier; Configure per-task resource floors and ceilings so eval scores reflect model capability rather than infrastructure; Read agentic coding leaderboards skeptically when gaps are under three points and configurations are undocumented; Account for confounders such as time-of-day API latency when comparing agent pass rates across runs; Design a controlled resource-sweep experiment holding model, harness and task set constant.
Topics
Sources
This page was written from 3 sources, 2 on domains other than anthropic.com.