VoiceAgentEval: A Dual-Dimensional Benchmark for Expert-Level Intelligent Voice-Agent Evaluation of Xbench's Professional-Aligned Series
Pengyu Xu, Shijia Li, Ao Sun, Feng Zhang, Yahan Li, Bo Wu, Zhanyu Ma, Jiguo Li, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, Rui Wang, Yang Liu, Xiaobo Hu, Fan Yang, Jia Zheng, Guanghua Yao
TL;DR
VoiceAgentEval addresses the lack of robust outbound-calling benchmarks by introducing a domain-rich, six-domain, 30-subscenario corpus and a large-language-model driven User Simulator to enable realistic, controllable testing. The framework combines a dual-layer textual evaluation (Task Flow Compliance and General Interaction Capability) with a comprehensive speech evaluation, anchored by 15 metrics across multiple domains and reinforced by human-in-the-loop verification. Experimental results across 12 state-of-the-art LLMs reveal trade-offs between task execution accuracy and interaction fluency, highlighting that model performance is shaped more by alignment and training data than sheer scale. This work provides a practical, extensible standard for benchmarking professional outbound AI systems, enabling targeted improvements in both dialogue generation and voice interaction quality.
Abstract
We propose OutboundEval, a comprehensive benchmark for evaluating large language models (LLMs) in expert-level intelligent outbound calling scenarios. Unlike existing methods that suffer from three key limitations - insufficient dataset diversity and category coverage, unrealistic user simulation, and inaccurate evaluation metrics - OutboundEval addresses these issues through a structured framework. First, we design a benchmark spanning six major business domains and 30 representative sub-scenarios, each with scenario-specific process decomposition, weighted scoring, and domain-adaptive metrics. Second, we develop a large-model-driven User Simulator that generates diverse, persona-rich virtual users with realistic behaviors, emotional variability, and communication styles, providing a controlled yet authentic testing environment. Third, we introduce a dynamic evaluation method that adapts to task variations, integrating automated and human-in-the-loop assessment to measure task execution accuracy, professional knowledge application, adaptability, and user experience quality. Experiments on 12 state-of-the-art LLMs reveal distinct trade-offs between expert-level task completion and interaction fluency, offering practical insights for building reliable, human-like outbound AI systems. OutboundEval establishes a practical, extensible, and domain-oriented standard for benchmarking LLMs in professional applications.
