Evaluation and Benchmarking of LLM Agents: A Survey

by Mohammadi, Li, Lo & Yip (KDD '25)

AdvancedPaperFree11 pages, ~1 hour read

A peer-reviewed map of every way people are currently measuring whether an agent works.

Start LearningAdded Jul 18, 2026 · Updated Aug 15, 2026

Overview

This eleven-page survey by Mahmoud Mohammadi, Yipeng Li, Jane Lo and Wendy Yip was posted to arXiv on 29 July 2025 and published at KDD '25 (3-7 August 2025) under DOI 10.1145/3711896.3736570. Its organising contribution is a two-dimensional taxonomy: one axis is evaluation objectives - what you are actually measuring, spanning agent behaviour, capabilities, reliability and safety - and the other is the evaluation process, covering interaction modes, datasets and benchmarks, metric-computation methods, and tooling. Against that frame it catalogues the benchmark landscape by domain: web interaction (WebArena, VisualWebArena, WebShop, BrowserGym, MiniWob), software engineering (SWE-bench, AutoCodeRover, AgentBoard), general and multi-task reasoning (AgentBench, GAIA, TaskBench, AppWorld), domain-specific suites (ScienceAgentBench, FinCon, ResearchArena, MobileAgentBench, GameBench), and a distinct safety and security cluster (AgentSecBench, SafeAgentBench, AgentPoison, CyberBench). The paper's sharpest argument is about what the field ignores: enterprise-specific requirements such as role-based access to data, hard reliability guarantees, dynamic long-horizon interactions and regulatory compliance are largely absent from published benchmarks, and it closes by calling for holistic, more realistic and scalable evaluation. With over a hundred references packed into eleven pages it functions best as a reading map and benchmark index rather than a standalone tutorial.

At a Glance

Topic
Agentic
Level
Advanced
Format
Paper
Cost
Free
Duration
11 pages, ~1 hour read
Provider
Mohammadi, Li, Lo & Yip (KDD '25)
Hands-on
No
Certificate
None

What You’ll Learn

  • Apply a two-dimensional taxonomy separating evaluation objectives from evaluation process
  • Distinguish measuring agent behaviour, capabilities, reliability and safety separately
  • Map benchmark suites to domains: web, code, science, finance and mobile
  • Choose interaction modes and metric-computation methods for a given agent
  • Locate the right benchmark from a catalogue of twenty-plus named suites
  • Recognise the safety and security benchmark cluster as a separate concern
  • Identify enterprise gaps: role-based data access, compliance, long-horizon interaction

Highlights

  • Peer-reviewed and published at KDD '25, not an arXiv-only preprint
  • Names enterprise constraints most academic evaluation surveys leave out entirely
  • Catalogues safety and security benchmarks alongside pure capability ones
  • Compact at eleven pages with 100+ references, usable as a citation index
  • Covers tooling and metric computation, not just the benchmark list

Who It’s For

Best For

  • Engineers choosing a benchmark suite for an agent already in production
  • Researchers who need a fast orientation in agent-evaluation literature
  • Platform teams designing internal agent reliability and safety metrics

Prerequisites

  • Familiarity with LLM agents, tool use and planning loops
  • Comfort reading an academic survey and chasing its citations

FAQ

What is Evaluation and Benchmarking of LLM Agents: A Survey?

A KDD '25 survey that organises the fragmented literature on LLM agent evaluation into a two-dimensional taxonomy of what to evaluate and how to evaluate it, then catalogues the named benchmark suites by domain. It is aimed at engineers and researchers who need to pick a benchmark for an agent they have already built, and at teams designing internal reliability metrics.

Is Evaluation and Benchmarking of LLM Agents: A Survey free?

Evaluation and Benchmarking of LLM Agents: A Survey is free to access.

What level is Evaluation and Benchmarking of LLM Agents: A Survey for?

Evaluation and Benchmarking of LLM Agents: A Survey is aimed at a advanced audience. Recommended background: Familiarity with LLM agents, tool use and planning loops, Comfort reading an academic survey and chasing its citations.

How long does Evaluation and Benchmarking of LLM Agents: A Survey take?

Expect roughly 11 pages, ~1 hour read. Most learners work through it at their own pace.

What will I learn from Evaluation and Benchmarking of LLM Agents: A Survey?

You'll learn: Apply a two-dimensional taxonomy separating evaluation objectives from evaluation process; Distinguish measuring agent behaviour, capabilities, reliability and safety separately; Map benchmark suites to domains: web, code, science, finance and mobile; Choose interaction modes and metric-computation methods for a given agent; Locate the right benchmark from a catalogue of twenty-plus named suites; Recognise the safety and security benchmark cluster as a separate concern; Identify enterprise gaps: role-based data access, compliance, long-horizon interaction.

Topics

llm-agentsevaluationbenchmarksagent-reliabilitysurvey

Sources

This page was written from 2 sources, 1 on domains other than arxiv.org.

  1. 1.arxiv.org2507.21504vendor
  2. 2.huggingface.co2507.21504