Table of Contents
Fetching ...

FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis

Fengbin Zhu, Xiang Yao Ng, Ziyang Liu, Chang Liu, Xianwei Zeng, Chao Wang, Tianhui Tan, Xuan Yao, Pengyang Shao, Min Xu, Zixuan Wang, Jing Wang, Xin Lin, Junfeng Li, Jingxian Zhu, Yang Zhang, Wenjie Wang, Fuli Feng, Richang Hong, Huanbo Luan, Ke-Wei Huang, Tat-Seng Chua

TL;DR

This work tackles the gap in rigorous evaluation of Deep Research (DR) agents for critical financial analysis by introducing HisRubric, a framework combining a hierarchical, expert-designed structure with a fine-grained rubric to assess four capabilities: Recognition, Calculation, Abstraction, and Interpretation. Building on this framework, the FinDeepResearch benchmark covers $64$ listed companies across $8$ markets and $4$ languages, totaling $15,808$ grading items across $6$ sections and $18$ subsections, enabling multi-faceted, verifiable assessment. The authors evaluate $16$ methods (6 DR agents, 5 thinking+search LLMs, 5 thinking-only LLMs), finding that DR agents generally outperform others and excel in Recognition and Calculation, while all approaches struggle with Interpretation and with analyses in non-English markets. The results highlight the practical importance of integrating rigorous structure with external retrieval for high-stakes financial research and point to future work in expanding multilingual capabilities and sharpening interpretive reasoning in DR systems.

Abstract

Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks. However, existing literature lacks a rigorous and systematic evaluation of DR Agent's capabilities in critical research analysis. To address this gap, we first propose HisRubric, a novel evaluation framework with a hierarchical analytical structure and a fine-grained grading rubric for rigorously assessing DR agents' capabilities in corporate financial analysis. This framework mirrors the professional analyst's workflow, progressing from data recognition to metric calculation, and finally to strategic summarization and interpretation. Built on this framework, we construct a FinDeepResearch benchmark that comprises 64 listed companies from 8 financial markets across 4 languages, encompassing a total of 15,808 grading items. We further conduct extensive experiments on the FinDeepResearch using 16 representative methods, including 6 DR agents, 5 LLMs equipped with both deep reasoning and search capabilities, and 5 LLMs with deep reasoning capabilities only. The results reveal the strengths and limitations of these approaches across diverse capabilities, financial markets, and languages, offering valuable insights for future research and development. The benchmark and evaluation code will be made publicly available.

FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis

TL;DR

This work tackles the gap in rigorous evaluation of Deep Research (DR) agents for critical financial analysis by introducing HisRubric, a framework combining a hierarchical, expert-designed structure with a fine-grained rubric to assess four capabilities: Recognition, Calculation, Abstraction, and Interpretation. Building on this framework, the FinDeepResearch benchmark covers listed companies across markets and languages, totaling grading items across sections and subsections, enabling multi-faceted, verifiable assessment. The authors evaluate methods (6 DR agents, 5 thinking+search LLMs, 5 thinking-only LLMs), finding that DR agents generally outperform others and excel in Recognition and Calculation, while all approaches struggle with Interpretation and with analyses in non-English markets. The results highlight the practical importance of integrating rigorous structure with external retrieval for high-stakes financial research and point to future work in expanding multilingual capabilities and sharpening interpretive reasoning in DR systems.

Abstract

Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks. However, existing literature lacks a rigorous and systematic evaluation of DR Agent's capabilities in critical research analysis. To address this gap, we first propose HisRubric, a novel evaluation framework with a hierarchical analytical structure and a fine-grained grading rubric for rigorously assessing DR agents' capabilities in corporate financial analysis. This framework mirrors the professional analyst's workflow, progressing from data recognition to metric calculation, and finally to strategic summarization and interpretation. Built on this framework, we construct a FinDeepResearch benchmark that comprises 64 listed companies from 8 financial markets across 4 languages, encompassing a total of 15,808 grading items. We further conduct extensive experiments on the FinDeepResearch using 16 representative methods, including 6 DR agents, 5 LLMs equipped with both deep reasoning and search capabilities, and 5 LLMs with deep reasoning capabilities only. The results reveal the strengths and limitations of these approaches across diverse capabilities, financial markets, and languages, offering valuable insights for future research and development. The benchmark and evaluation code will be made publicly available.
Paper Structure (24 sections, 1 equation, 9 figures, 7 tables)

This paper contains 24 sections, 1 equation, 9 figures, 7 tables.

Figures (9)

  • Figure 1: An overview of the HisRubric evaluation framework.The numbers in brackets indicate the number of grading items (left) and the corresponding full marks (right).
  • Figure 2: An overview for constructing FinDeepResearch.
  • Figure 3: An evaluation of representative methods on FinDeepResearch w.r.t Information Precision.
  • Figure 4: An evaluation of representative methods on FinDeepResearch w.r.t Structural Rigor.
  • Figure 5: Performance analysis across four different capabilities.
  • ...and 4 more figures