Subgroup Performance Analysis in Hidden Stratifications

Alceu Bissoto; Trung-Dung Hoang; Tim Flühmann; Susu Sun; Christian F. Baumgartner; Lisa M. Koch

Subgroup Performance Analysis in Hidden Stratifications

Alceu Bissoto, Trung-Dung Hoang, Tim Flühmann, Susu Sun, Christian F. Baumgartner, Lisa M. Koch

TL;DR

This work tackles the challenge of hidden stratifications causing performance disparities in medical imaging models by introducing subgroup discovery as a tool for performance monitoring. The authors implement a DOMINO-based approach that clusters task-agnostic image representations with a balancing parameter $ abla$, enabling discovery of cohesive subgroups without requiring ground-truth subgroup labels. Through synthetic artifacts and real-world datasets (CheXpert-Plus and SLICE-3D), they show that discovered subgroups reveal larger performance gaps than metadata-based subgroups while maintaining cohesion, and that these gaps align with meaningful visual features rather than demographics. Their results justify using subgroup discovery as a practical, robust complement to traditional subgroup analysis for safer and more trustworthy AI deployment in medicine, with implications for both validation and bias mitigation. Key metrics include the performance gap $Δ(S)$ and average purity $AP(S)$, and the findings demonstrate that natural image feature extractors are often sufficient to expose important disparities, supporting broader applicability in clinical settings.

Abstract

Machine learning (ML) models may suffer from significant performance disparities between patient groups. Identifying such disparities by monitoring performance at a granular level is crucial for safely deploying ML to each patient. Traditional subgroup analysis based on metadata can expose performance disparities only if the available metadata (e.g., patient sex) sufficiently reflects the main reasons for performance variability, which is not common. Subgroup discovery techniques that identify cohesive subgroups based on learned feature representations appear as a potential solution: They could expose hidden stratifications and provide more granular subgroup performance reports. However, subgroup discovery is challenging to evaluate even as a standalone task, as ground truth stratification labels do not exist in real data. Subgroup discovery has thus neither been applied nor evaluated for the application of subgroup performance monitoring. Here, we apply subgroup discovery for performance monitoring in chest x-ray and skin lesion classification. We propose novel evaluation strategies and show that a simplified subgroup discovery method without access to classification labels or metadata can expose larger performance disparities than traditional metadata-based subgroup analysis. We provide the first compelling evidence that subgroup discovery can serve as an important tool for comprehensive performance validation and monitoring of trustworthy AI in medicine.

Subgroup Performance Analysis in Hidden Stratifications

TL;DR

Abstract

Subgroup Performance Analysis in Hidden Stratifications

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (4)