Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness

Khyathi Raghavi Chandu; Linjie Li; Anas Awadalla; Ximing Lu; Jae Sung Park; Jack Hessel; Lijuan Wang; Yejin Choi

Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness

Khyathi Raghavi Chandu, Linjie Li, Anas Awadalla, Ximing Lu, Jae Sung Park, Jack Hessel, Lijuan Wang, Yejin Choi

TL;DR

This work addresses the scarcity of reliable uncertainty handling in vision-language models by introducing a taxonomy that separates epistemic and aleatoric uncertainty and their fine-grained subcategories. It builds the CertainlyUncertain dataset (~178K VQA samples) through two data-generation pipelines: image perturbations to create unanswerable questions and caption-driven QA generation, yielding rich, contrastive pairs. A new confidence-weighted accuracy metric is proposed to jointly capture correctness and calibration, showing stronger alignment with accuracy and lower calibration error than prior metrics. Experiments across multiple base models and training strategies demonstrate improved refusal behavior and reduced hallucinations, while preserving standard VQA performance, highlighting the practical value of structured uncertainty awareness for robust AI systems.

Abstract

The ability to acknowledge the inevitable uncertainty in their knowledge and reasoning is a prerequisite for AI systems to be truly truthful and reliable. In this paper, we present a taxonomy of uncertainty specific to vision-language AI systems, distinguishing between epistemic uncertainty (arising from a lack of information) and aleatoric uncertainty (due to inherent unpredictability), and further explore finer categories within. Based on this taxonomy, we synthesize a benchmark dataset, CertainlyUncertain, featuring 178K visual question answering (VQA) samples as contrastive pairs. This is achieved by 1) inpainting images to make previously answerable questions into unanswerable ones; and 2) using image captions to prompt large language models for both answerable and unanswerable questions. Additionally, we introduce a new metric confidence-weighted accuracy, that is well correlated with both accuracy and calibration error, to address the shortcomings of existing metrics.

Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness

TL;DR

Abstract

Paper Structure (21 sections, 2 equations, 11 figures, 8 tables)

This paper contains 21 sections, 2 equations, 11 figures, 8 tables.

Introduction
CertainlyUncertain
Taxonomy of Uncertainty Awareness
Dataset Creation
Evaluation Metrics
Experiments
Experimental Details
Evaluation Benchmarks
Results and Discussion
Related Work
Conclusions and Future Work
Limitations
Broader Impact
Samples visualizing CertainlyUncertain benchmark
Samples visualizing predictions and confidence-weighted metric
...and 6 more sections

Figures (11)

Figure 1: CertainlyUncertain: Taxonomy of uncertainty awareness in multimodal reasoning
Figure 2: Pipeline for sourcing from images
Figure 3: Uncertainty paradox in generative VLMs, where the question is generated from GPT-4/GPT-4V.
Figure 4: Correlation of confidence weighted accuracy ($\uparrow$) with $\text{LAVE}_{\text{idk}}$ accuracy ($\uparrow$) and ECE ($\downarrow$). The datapoints in this plot are from evaluation results on extraneous split of different model variants in our experiments.
Figure 5: Breakdown of model performance on finegrained categories. We report $\text{LAVE}_{\text{idk}}$ Metric Accuracy as the confidence of GPT-4V prediction is not accessible.
...and 6 more figures

Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness

TL;DR

Abstract

Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness

Authors

TL;DR

Abstract

Table of Contents

Figures (11)