Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

Sayash Kapoor; Benedikt Stroebl; Peter Kirgis; Nitya Nadgir; Zachary S Siegel; Boyi Wei; Tianci Xue; Ziru Chen; Felix Chen; Saiteja Utpala; Franck Ndzomga; Dheeraj Oruganty; Sophie Luskin; Kangheng Liu; Botao Yu; Amit Arora; Dongyoon Hahm; Harsh Trivedi; Huan Sun; Juyong Lee; Tengjun Jin; Yifan Mai; Yifei Zhou; Yuxuan Zhu; Rishi Bommasani; Daniel Kang; Dawn Song; Peter Henderson; Yu Su; Percy Liang; Arvind Narayanan

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

Sayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir, Zachary S Siegel, Boyi Wei, Tianci Xue, Ziru Chen, Felix Chen, Saiteja Utpala, Franck Ndzomga, Dheeraj Oruganty, Sophie Luskin, Kangheng Liu, Botao Yu, Amit Arora, Dongyoon Hahm, Harsh Trivedi, Huan Sun, Juyong Lee, Tengjun Jin, Yifan Mai, Yifei Zhou, Yuxuan Zhu, Rishi Bommasani, Daniel Kang, Dawn Song, Peter Henderson, Yu Su, Percy Liang, Arvind Narayanan

TL;DR

The paper tackles the fragmented and costly evaluation of AI agents across real-world domains. It introduces the Holistic Agent Leaderboard (HAL) harness, a unified, scalable framework for running and logging agent evaluations on many benchmarks and models, with automated log analysis to detect bugs and unsafe behaviors. Key findings include that higher reasoning effort often does not improve accuracy, that agent scaffolds profoundly affect cost and performance, and that log-based analysis reveals issues like shortcuts and data leakage that pure accuracy metrics miss. HAL enables reproducible, cost-aware, cross-domain benchmarking and provides a data-rich resource (2.5B tokens) to study agent behavior, with an aim to shift focus toward reliable real-world performance.

Abstract

AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of how well agents really work. We introduce the Holistic Agent Leaderboard (HAL) to address these challenges. We make three main contributions. First, we provide a standardized evaluation harness that orchestrates parallel evaluations across hundreds of VMs, reducing evaluation time from weeks to hours while eliminating common implementation bugs. Second, we conduct three-dimensional analysis spanning models, scaffolds, and benchmarks. We validate the harness by conducting 21,730 agent rollouts across 9 models and 9 benchmarks in coding, web navigation, science, and customer service with a total cost of about $40,000. Our analysis reveals surprising insights, such as higher reasoning effort reducing accuracy in the majority of runs. Third, we use LLM-aided log inspection to uncover previously unreported behaviors, such as searching for the benchmark on HuggingFace instead of solving a task, or misusing credit cards in flight booking tasks. We share all agent logs, comprising 2.5B tokens of language model calls, to incentivize further research into agent behavior. By standardizing how the field evaluates agents and addressing common pitfalls in agent evaluation, we hope to shift the focus from agents that ace benchmarks to agents that work reliably in the real world.

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

TL;DR

Abstract

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (48)