Table of Contents
Fetching ...

FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain

Tiansheng Hu, Tongyan Hu, Liuyang Bai, Yilun Zhao, Arman Cohan, Chen Zhao

TL;DR

FinTrust presents a holistic benchmark for evaluating trustworthiness of finance-oriented LLMs across seven dimensions (Truthfulness, Safety, Fairness, Robustness, Privacy, Transparency, Knowledge Discovery) organized into three pragmatic subsets with multimodal data (text, tables, time-series). It introduces fine-grained tasks and real-world scenarios, comprising $15,680$ QA instances evaluated on $11$ LLMs (proprietary, open-source, and finance-domain specific). Key findings show proprietary models generally excel in safety and trustworthiness, open-source models demonstrate strengths in industry fairness, and all models exhibit gaps in fiduciary alignment and disclosure, highlighting substantial room for improvement. FinTrust thus provides actionable insights for deploying LLMs in high-stakes finance and outlines concrete directions for future alignment, safety, and transparency improvements in domain-specific contexts.

Abstract

Recent LLMs have demonstrated promising ability in solving finance related problems. However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property. This paper introduces FinTrust, a comprehensive benchmark specifically designed for evaluating the trustworthiness of LLMs in finance applications. Our benchmark focuses on a wide range of alignment issues based on practical context and features fine-grained tasks for each dimension of trustworthiness evaluation. We assess eleven LLMs on FinTrust and find that proprietary models like o4-mini outperforms in most tasks such as safety while open-source models like DeepSeek-V3 have advantage in specific areas like industry-level fairness. For challenging task like fiduciary alignment and disclosure, all LLMs fall short, showing a significant gap in legal awareness. We believe that FinTrust can be a valuable benchmark for LLMs' trustworthiness evaluation in finance domain.

FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain

TL;DR

FinTrust presents a holistic benchmark for evaluating trustworthiness of finance-oriented LLMs across seven dimensions (Truthfulness, Safety, Fairness, Robustness, Privacy, Transparency, Knowledge Discovery) organized into three pragmatic subsets with multimodal data (text, tables, time-series). It introduces fine-grained tasks and real-world scenarios, comprising QA instances evaluated on LLMs (proprietary, open-source, and finance-domain specific). Key findings show proprietary models generally excel in safety and trustworthiness, open-source models demonstrate strengths in industry fairness, and all models exhibit gaps in fiduciary alignment and disclosure, highlighting substantial room for improvement. FinTrust thus provides actionable insights for deploying LLMs in high-stakes finance and outlines concrete directions for future alignment, safety, and transparency improvements in domain-specific contexts.

Abstract

Recent LLMs have demonstrated promising ability in solving finance related problems. However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property. This paper introduces FinTrust, a comprehensive benchmark specifically designed for evaluating the trustworthiness of LLMs in finance applications. Our benchmark focuses on a wide range of alignment issues based on practical context and features fine-grained tasks for each dimension of trustworthiness evaluation. We assess eleven LLMs on FinTrust and find that proprietary models like o4-mini outperforms in most tasks such as safety while open-source models like DeepSeek-V3 have advantage in specific areas like industry-level fairness. For challenging task like fiduciary alignment and disclosure, all LLMs fall short, showing a significant gap in legal awareness. We believe that FinTrust can be a valuable benchmark for LLMs' trustworthiness evaluation in finance domain.
Paper Structure (58 sections, 1 equation, 36 figures, 6 tables)

This paper contains 58 sections, 1 equation, 36 figures, 6 tables.

Figures (36)

  • Figure 1: An overview of the seven dimensions of trustworthiness assessed in the FinTrust benchmark. FinTrust distinguishes from existing benchmarks featuring three unique characteristics: (1) Alignment Evaluation: The Safety, Fairness, Privacy and Transparency dimensions are specifically designed to assess the legal and ethical aspects of trustworthiness; (2) Fine-Grained Tasks: We design multiple tasks under each dimension. In particular, Trustfulness deals with both hallucination and number calculation; Safety includes four different attack methods; Fairness evaluation covers both industry-level and personal-level; Privacy features three types of system prompts with different levels of emphasis on privacy awareness; (3) Real-World Scenarios: We imitate the challenges that are from real applications. For example, Safety evaluation includes ten financial crimes.
  • Figure 2: Safety evaluation with LLM-as-a-judge. Genetic Algorithm attack is the only effective attack to most LLMs except o4-mini.
  • Figure 3: Personal Level Fairness Analysis. Fin-R1 outperforms all the other models in correctness and stability while Llama 4 is unstable to sensitive attribution changes.
  • Figure 4: Privacy Analysis with LLM-as-a-judge under different system prompts on privacy issues (not mention, implicit mention and explicit mention). o4-mini demonstrates the best privacy alertness. All the finance domain specific LLMs are weak in this category.
  • Figure 5: Transparency Analysis Across LLMs. Ideally, LLMs should consistently select Company A without referencing ownership information. However, most models tend to favor Company B when the system prompt specifies ownership as Company B.
  • ...and 31 more figures