Evaluation and Benchmarking of LLM Agents: A Survey
by Mohammadi, Li, Lo & Yip (KDD '25)
A peer-reviewed map of every way people are currently measuring whether an agent works.
Overview
This eleven-page survey by Mahmoud Mohammadi, Yipeng Li, Jane Lo and Wendy Yip was posted to arXiv on 29 July 2025 and published at KDD '25 (3-7 August 2025) under DOI 10.1145/3711896.3736570. Its organising contribution is a two-dimensional taxonomy: one axis is evaluation objectives - what you are actually measuring, spanning agent behaviour, capabilities, reliability and safety - and the other is the evaluation process, covering interaction modes, datasets and benchmarks, metric-computation methods, and tooling. Against that frame it catalogues the benchmark landscape by domain: web interaction (WebArena, VisualWebArena, WebShop, BrowserGym, MiniWob), software engineering (SWE-bench, AutoCodeRover, AgentBoard), general and multi-task reasoning (AgentBench, GAIA, TaskBench, AppWorld), domain-specific suites (ScienceAgentBench, FinCon, ResearchArena, MobileAgentBench, GameBench), and a distinct safety and security cluster (AgentSecBench, SafeAgentBench, AgentPoison, CyberBench). The paper's sharpest argument is about what the field ignores: enterprise-specific requirements such as role-based access to data, hard reliability guarantees, dynamic long-horizon interactions and regulatory compliance are largely absent from published benchmarks, and it closes by calling for holistic, more realistic and scalable evaluation. With over a hundred references packed into eleven pages it functions best as a reading map and benchmark index rather than a standalone tutorial.
At a Glance
- Topic
- Agentic
- Level
- Advanced
- Format
- Paper
- Cost
- Free
- Duration
- 11 pages, ~1 hour read
- Provider
- Mohammadi, Li, Lo & Yip (KDD '25)
- Hands-on
- No
- Certificate
- None
What You’ll Learn
- ✓Apply a two-dimensional taxonomy separating evaluation objectives from evaluation process
- ✓Distinguish measuring agent behaviour, capabilities, reliability and safety separately
- ✓Map benchmark suites to domains: web, code, science, finance and mobile
- ✓Choose interaction modes and metric-computation methods for a given agent
- ✓Locate the right benchmark from a catalogue of twenty-plus named suites
- ✓Recognise the safety and security benchmark cluster as a separate concern
- ✓Identify enterprise gaps: role-based data access, compliance, long-horizon interaction
Highlights
- •Peer-reviewed and published at KDD '25, not an arXiv-only preprint
- •Names enterprise constraints most academic evaluation surveys leave out entirely
- •Catalogues safety and security benchmarks alongside pure capability ones
- •Compact at eleven pages with 100+ references, usable as a citation index
- •Covers tooling and metric computation, not just the benchmark list
Who It’s For
Best For
- ✓Engineers choosing a benchmark suite for an agent already in production
- ✓Researchers who need a fast orientation in agent-evaluation literature
- ✓Platform teams designing internal agent reliability and safety metrics
Prerequisites
- •Familiarity with LLM agents, tool use and planning loops
- •Comfort reading an academic survey and chasing its citations
FAQ
What is Evaluation and Benchmarking of LLM Agents: A Survey?
A KDD '25 survey that organises the fragmented literature on LLM agent evaluation into a two-dimensional taxonomy of what to evaluate and how to evaluate it, then catalogues the named benchmark suites by domain. It is aimed at engineers and researchers who need to pick a benchmark for an agent they have already built, and at teams designing internal reliability metrics.
Is Evaluation and Benchmarking of LLM Agents: A Survey free?
Evaluation and Benchmarking of LLM Agents: A Survey is free to access.
What level is Evaluation and Benchmarking of LLM Agents: A Survey for?
Evaluation and Benchmarking of LLM Agents: A Survey is aimed at a advanced audience. Recommended background: Familiarity with LLM agents, tool use and planning loops, Comfort reading an academic survey and chasing its citations.
How long does Evaluation and Benchmarking of LLM Agents: A Survey take?
Expect roughly 11 pages, ~1 hour read. Most learners work through it at their own pace.
What will I learn from Evaluation and Benchmarking of LLM Agents: A Survey?
You'll learn: Apply a two-dimensional taxonomy separating evaluation objectives from evaluation process; Distinguish measuring agent behaviour, capabilities, reliability and safety separately; Map benchmark suites to domains: web, code, science, finance and mobile; Choose interaction modes and metric-computation methods for a given agent; Locate the right benchmark from a catalogue of twenty-plus named suites; Recognise the safety and security benchmark cluster as a separate concern; Identify enterprise gaps: role-based data access, compliance, long-horizon interaction.
Topics
Sources
This page was written from 2 sources, 1 on domains other than arxiv.org.