The challenge of uncertainty quantification of large language models in medicine

Zahra Atf; Seyed Amir Ahmad Safavi-Naini; Peter R. Lewis; Aref Mahjoubfar; Nariman Naderi; Thomas R. Savage; Ali Soroush

The challenge of uncertainty quantification of large language models in medicine

Zahra Atf, Seyed Amir Ahmad Safavi-Naini, Peter R. Lewis, Aref Mahjoubfar, Nariman Naderi, Thomas R. Savage, Ali Soroush

TL;DR

This paper addresses the critical challenge of uncertainty quantification in large language models (LLMs) for medical applications, arguing that communicating uncertainty is essential for safe, trustworthy AI-assisted care. It proposes a comprehensive framework that fuses Bayesian inference, deep ensembles, Monte Carlo dropout, and linguistic entropy with surrogate modeling, multi-source data integration, dynamic calibration, continual/meta-learning, and explainability via uncertainty maps and confidence metrics. Emphasizing Responsible and Reflective AI, the work integrates technical and philosophical perspectives to promote transparency, accountability, and ethical deployment in high-stakes clinical settings. The proposed approach aims to improve trust, safety, and interpretability of AI-driven medical decisions while accommodating proprietary API limitations and evolving medical knowledge.

Abstract

This study investigates uncertainty quantification in large language models (LLMs) for medical applications, emphasizing both technical innovations and philosophical implications. As LLMs become integral to clinical decision-making, accurately communicating uncertainty is crucial for ensuring reliable, safe, and ethical AI-assisted healthcare. Our research frames uncertainty not as a barrier but as an essential part of knowledge that invites a dynamic and reflective approach to AI design. By integrating advanced probabilistic methods such as Bayesian inference, deep ensembles, and Monte Carlo dropout with linguistic analysis that computes predictive and semantic entropy, we propose a comprehensive framework that manages both epistemic and aleatoric uncertainties. The framework incorporates surrogate modeling to address limitations of proprietary APIs, multi-source data integration for better context, and dynamic calibration via continual and meta-learning. Explainability is embedded through uncertainty maps and confidence metrics to support user trust and clinical interpretability. Our approach supports transparent and ethical decision-making aligned with Responsible and Reflective AI principles. Philosophically, we advocate accepting controlled ambiguity instead of striving for absolute predictability, recognizing the inherent provisionality of medical knowledge.

The challenge of uncertainty quantification of large language models in medicine

TL;DR

Abstract

The challenge of uncertainty quantification of large language models in medicine

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (11)