AgenticMCPFrameworks

Agent Harness Engineering: A Survey

by Junjie Li, Chandan K Reddy et al. (CMU, Virginia Tech, Amazon, Stanford and others)

IntermediatePaperFreeSurvey paper plus a 372-entry companion list, self-paced reading

A seven-layer map (ETCLOVG) of the infrastructure around an LLM agent, with 138 open-source projects sorted onto it.

Start LearningAdded Oct 2, 2026 · Updated Oct 2, 2026

Overview

Agent Harness Engineering: A Survey, posted on OpenReview in May 2026 by Junjie Li, Xi Xiao, Chandan K Reddy and co-authors from Carnegie Mellon, Stanford, Yale, Virginia Tech, Amazon, the University of Chicago and other institutions, treats the agent execution harness as its own system layer. Its central claim is that task reliability in production agents depends less on the base model than on the infrastructure wrapped around it. The paper traces a progression from prompt engineering to context engineering to harness engineering, then proposes ETCLOVG, a seven-layer taxonomy: Execution environment and sandbox, Tool interface and protocol, Context and memory management, Lifecycle and orchestration, Observability and operations, Verification and evaluation, and Governance and security. Observability and governance get their own layers because in production each has its own tooling and usually its own owning team. The authors map 138 open-source projects onto the layers (47 in orchestration, 21 in evaluation, 20 in execution, 15 in observability, 14 in governance, 12 in tools, 9 in context), describe a five-stage evaluation lifecycle that starts with readiness validation, and name an observability-evaluation gap: teams log traces but rarely turn them into offline evals or regression tests. A companion GitHub list, Picrew/awesome-agent-harness, catalogs 372 resources and was last verified September 21, 2026.

At a Glance

Topic
Agentic
Level
Intermediate
Format
Paper
Cost
Free
Duration
Survey paper plus a 372-entry companion list, self-paced reading
Provider
Junjie Li, Chandan K Reddy et al. (CMU, Virginia Tech, Amazon, Stanford and others)
Hands-on
No
Certificate
None

What You’ll Learn

  • ✓Classify any agent stack against the seven ETCLOVG harness layers to find missing components
  • ✓Separate observability and governance concerns from orchestration when designing an agent platform
  • ✓Recognize the observability-evaluation gap and convert production traces into regression tests
  • ✓Apply a five-stage evaluation lifecycle that begins with readiness validation before execution
  • ✓Reason about coupling effects where a change in one harness layer shifts measurements elsewhere
  • ✓Pick open-source sandboxes, MCP servers, memory layers and eval tools from a layered catalog

Highlights

  • •Treats observability and governance as independent layers, which earlier six-component harness frameworks folded into other parts
  • •Grounds the taxonomy in 138 mapped open-source projects, with per-layer counts that show where the ecosystem is thin (only 9 in context management)
  • •Companion awesome-agent-harness repo has 1,821 stars and 211 forks and was pushed September 20, 2026 (checked October 2, 2026)
  • •An independent review on ai-eval.org notes its limits honestly: it is a taxonomy and synthesis, not a controlled benchmark, and some corpus counts differ between abstract, body and project page

Who It’s For

Best For

  • ✓Platform engineers designing or auditing the infrastructure around production LLM agents
  • ✓Teams choosing sandbox, tracing, eval and guardrail tools for an agent stack
  • ✓Engineering leads who need a shared vocabulary for agent reliability work
  • ✓Researchers looking for a structured map of agent harness literature and tooling

Prerequisites

  • •Working familiarity with LLM agents, tool calling and MCP
  • •Some exposure to production concerns such as tracing, evals and access control
  • •Comfort reading academic survey papers

FAQ

What is Agent Harness Engineering: A Survey?

Agent Harness Engineering: A Survey is a 2026 paper for engineers who run LLM agents in production. It argues that reliability depends more on the execution harness around the model than on the model itself, and gives you a seven-layer taxonomy for auditing sandboxes, tools, context, orchestration, observability, evaluation and governance in your own stack.

Is Agent Harness Engineering: A Survey free?

Agent Harness Engineering: A Survey is free to access.

What level is Agent Harness Engineering: A Survey for?

Agent Harness Engineering: A Survey is aimed at a intermediate audience. Recommended background: Working familiarity with LLM agents, tool calling and MCP, Some exposure to production concerns such as tracing, evals and access control, Comfort reading academic survey papers.

How long does Agent Harness Engineering: A Survey take?

Expect roughly Survey paper plus a 372-entry companion list, self-paced reading. Most learners work through it at their own pace.

What will I learn from Agent Harness Engineering: A Survey?

You'll learn: Classify any agent stack against the seven ETCLOVG harness layers to find missing components; Separate observability and governance concerns from orchestration when designing an agent platform; Recognize the observability-evaluation gap and convert production traces into regression tests; Apply a five-stage evaluation lifecycle that begins with readiness validation before execution; Reason about coupling effects where a change in one harness layer shifts measurements elsewhere; Pick open-source sandboxes, MCP servers, memory layers and eval tools from a layered catalog.

Topics

agent harnessagent infrastructurellm agentsagent evaluationagent observabilitysurvey

Sources

This page was written from 4 sources, 3 on domains other than picrew.github.io.

  1. 1.picrew.github.io — LLM Harnessvendor
  2. 2.github.com — awesome agent harness
  3. 3.api.github.com — awesome agent harness
  4. 4.ai-eval.org — openreview agent harness engineering survey