LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News

Yunfan Zhang; Kathleen McKeown; Smaranda Muresan

LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News

Yunfan Zhang, Kathleen McKeown, Smaranda Muresan

Abstract

Large Language Models (LLMs) with agentic web search capabilities show strong potential for tasks requiring real-time information access and complex fact retrieval, yet evaluating such systems remains challenging. We introduce \bench, a rigorous and regularly updated benchmark designed to assess the agentic web search abilities of LLMs. \bench automatically generates fresh question-answer pairs from recent news articles, ensuring that questions require information beyond an LLM's training data and enabling clear separation between internal knowledge and search capability. The benchmark features intentionally difficult questions requiring multi-hop search queries, page visits, and reasoning, making it well-suited for evaluating agentic search behavior. Our automated data curation and question generation pipeline enables frequent benchmark updates and supports construction of a large-scale training dataset for agentic web search models, addressing the scarcity of such data in the research community. To ensure reliable evaluation, we include a subset of human-verified samples in the test set. We evaluate a broad range of systems using \bench, including commercial and open-weight LLMs as well as LLM-based web search APIs. The leaderboard, datasets, and code are publicly available at livenewsbench.com.

LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News

Abstract

Paper Structure (29 sections, 1 figure, 7 tables)

This paper contains 29 sections, 1 figure, 7 tables.

Introduction
Related Work
Regularly Updated "Live" Benchmarks for LLMs
Benchmarks for Agentic Search Evaluation
Dataset Construction
Retrieving News Articles
Dataset Partitioning
Experiment Setup
Agentic Web Search Framework
Evaluating Integrated LLM Search Systems
Judging Search Outputs
Results and Analysis
Comparison with other Time-Sensitive Factual QA Benchmarks
LiveNewsBench Human-Verified Test Set Results
Human-Verified Test Set vs. Full Test Set
...and 14 more sections

Figures (1)

Figure 1: Our automated dataset construction and one example from the Human-Verified Test Set. Our dataset construction pipeline comprises two main components: (1) retrieving news articles from online sources and (2) generating Q&A pairs from the retrieved content. Questions are designed to be challenging, requiring multiple searches, page visits, and reasoning steps.

LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News

Abstract

LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News

Authors

Abstract

Table of Contents

Figures (1)