An Empirical Study of Harness Design for Coding Agents
by Fan, Zhang, Ma et al. (UMass Amherst, Zoom, Emory, UNC Charlotte)
A 176-setting ablation of planning, tool sets and context management in coding agents, with numbers for when each one pays off.
Overview
This 43-page paper by Run-Ze Fan, Zihao Zhang, Simin Ma, Yebowen Hu, Shouju Wang, Kaiqiang Song, Fei Liu, Hamed Zamani and Xiaoyang Wang (UMass Amherst, Emory University and UNC Charlotte; the first authors did the work as interns at Zoom) was posted to arXiv on September 17, 2026. Most earlier work treats a coding harness as one unit. This study keeps the agent loop fixed and varies three components separately. Section 2 defines them: planning, the action space (a full predefined tool set versus bash only) and context management, with five strategies running from none (T0) through rule-based elision (T1), elision with a recall tool (T2) and LLM summarization (T3) to all three staged (T4). Section 3 runs 176 matched settings with Nemotron-3 30B, 120B and 550B and Mistral-Medium-3.5-128B on SWE-Bench Verified and Terminal-Bench 2.1 at 32k, 64k, 96k and 128k windows. Section 4 analyses trajectories. Context management closes a 35.7-point SWE-Bench gap at 32k to 2.7 points at 128k, mostly by preventing overflow, and staged T4 is cheapest at every budget. The recall tool added nothing (minus 0.36 points), and 56.3% of runs never called it. Planning adds 11.6 points for the 30B model but mainly cuts cost by about 30% for the strong models. Bash-only lifts the 550B model 3.6 points and cuts its cost 53%. The appendices publish the harness prompts and tool descriptions.
At a Glance
- Topic
- Agentic
- Level
- Intermediate
- Format
- Paper
- Cost
- Free
- Duration
- 43 pages, ~2-3 hours to read closely
- Provider
- Fan, Zhang, Ma et al. (UMass Amherst, Zoom, Emory, UNC Charlotte)
- Hands-on
- No
- Certificate
- None
What You’ll Learn
- ✓How context management effects shrink as the context window grows from 32k to 128k tokens
- ✓Why rule-based elision before LLM summarization gives the lowest cost per task at every budget
- ✓Why a recall tool for elided content added complexity that models rarely used
- ✓When an explicit planning step raises accuracy for weaker models and mainly cuts cost for stronger ones
- ✓When a bash-only action space beats a predefined tool set on cost and success rate
- ✓How to read agent trajectories to see whether a component changes behaviour or only run length
- ✓How to set up a controlled ablation of harness components on SWE-Bench Verified and Terminal-Bench
Highlights
- •Controlled ablations across 176 matched settings, which few harness write-ups attempt; most compare whole products
- •Concrete cost and accuracy deltas per component, such as a 53% cost cut from bash-only on the 550B model
- •Appendices publish the harness prompts, tool descriptions and trajectory analysis, so the setup can be reused
- •Drew 225 points and 59 comments on Hacker News, including praise from the mini-swe-agent author and the fair criticism that the models tested are not current frontier models
Who It’s For
Best For
- ✓Engineers building or tuning coding-agent harnesses
- ✓Teams deciding whether to ship a planning tool, a custom tool set or plain bash
- ✓Researchers designing ablations for agent scaffolding
- ✓Anyone running open-weight models under tight context budgets
Prerequisites
- •Familiarity with how coding agents loop over tool calls
- •Working knowledge of SWE-Bench-style evaluation and context windows
- •Comfort reading ablation tables in an ML paper
FAQ
What is An Empirical Study of Harness Design for Coding Agents?
An Empirical Study of Harness Design for Coding Agents is a September 2026 paper for engineers who build or tune coding-agent scaffolding. It holds the execution loop fixed and ablates planning, action space and context management across four models on SWE-Bench Verified and Terminal-Bench 2.1, so you can decide which harness components your model and context budget actually need.
Is An Empirical Study of Harness Design for Coding Agents free?
An Empirical Study of Harness Design for Coding Agents is free to access.
What level is An Empirical Study of Harness Design for Coding Agents for?
An Empirical Study of Harness Design for Coding Agents is aimed at a intermediate audience. Recommended background: Familiarity with how coding agents loop over tool calls, Working knowledge of SWE-Bench-style evaluation and context windows, Comfort reading ablation tables in an ML paper.
How long does An Empirical Study of Harness Design for Coding Agents take?
Expect roughly 43 pages, ~2-3 hours to read closely. Most learners work through it at their own pace.
What will I learn from An Empirical Study of Harness Design for Coding Agents?
You'll learn: How context management effects shrink as the context window grows from 32k to 128k tokens; Why rule-based elision before LLM summarization gives the lowest cost per task at every budget; Why a recall tool for elided content added complexity that models rarely used; When an explicit planning step raises accuracy for weaker models and mainly cuts cost for stronger ones; When a bash-only action space beats a predefined tool set on cost and success rate; How to read agent trajectories to see whether a component changes behaviour or only run length; How to set up a controlled ablation of harness components on SWE-Bench Verified and Terminal-Bench.
Topics
Sources
This page was written from 3 sources, 1 on domains other than arxiv.org.
- 1.arxiv.org — 2609.20804vendor
- 2.arxiv.org — 2609.20804v1vendor
- 3.hn.algolia.com — 49753878