nnActive: A Framework for Evaluation of Active Learning in 3D Biomedical Segmentation

Carsten T. Lüth; Jeremias Traub; Kim-Celine Kahl; Till J. Bungert; Lukas Klein; Lars Krämer; Paul F. Jaeger; Fabian Isensee; Klaus Maier-Hein

nnActive: A Framework for Evaluation of Active Learning in 3D Biomedical Segmentation

Carsten T. Lüth, Jeremias Traub, Kim-Celine Kahl, Till J. Bungert, Lukas Klein, Lars Krämer, Paul F. Jaeger, Fabian Isensee, Klaus Maier-Hein

TL;DR

nnActive delivers a rigorous, open-source framework to evaluate active learning for 3D biomedical segmentation across multiple datasets and budget regimes, addressing prior methodological pitfalls. It integrates partial-annotation training on 3D patches within an enhanced nnU-Net, introduces Foreground Aware Random baselines, and proposes the FG-Eff metric to better capture annotation effort. The large-scale study reveals that while AL methods outperform naive Random sampling, foreground-aware random baselines often challenge AL, with Predictive Entropy being a strong but variable performer. The work provides practical guidelines and a robust benchmark to catalyze further research and application of AL in 3D biomedical imaging.

Abstract

Semantic segmentation is crucial for various biomedical applications, yet its reliance on large annotated datasets presents a bottleneck due to the high cost and specialized expertise required for manual labeling. Active Learning (AL) aims to mitigate this challenge by querying only the most informative samples, thereby reducing annotation effort. However, in the domain of 3D biomedical imaging, there is no consensus on whether AL consistently outperforms Random sampling. Four evaluation pitfalls hinder the current methodological assessment. These are (1) restriction to too few datasets and annotation budgets, (2) using 2D models on 3D images without partial annotations, (3) Random baseline not being adapted to the task, and (4) measuring annotation cost only in voxels. In this work, we introduce nnActive, an open-source AL framework that overcomes these pitfalls by (1) means of a large scale study spanning four biomedical imaging datasets and three label regimes, (2) extending nnU-Net by using partial annotations for training with 3D patch-based query selection, (3) proposing Foreground Aware Random sampling strategies tackling the foreground-background class imbalance of medical images and (4) propose the foreground efficiency metric, which captures the low annotation cost of background-regions. We reveal the following findings: (A) while all AL methods outperform standard Random sampling, none reliably surpasses an improved Foreground Aware Random sampling; (B) benefits of AL depend on task specific parameters; (C) Predictive Entropy is overall the best performing AL method, but likely requires the most annotation effort; (D) AL performance can be improved with more compute intensive design choices. As a holistic, open-source framework, nnActive can serve as a catalyst for research and application of AL in 3D biomedical imaging. Code is at: https://github.com/MIC-DKFZ/nnActive

nnActive: A Framework for Evaluation of Active Learning in 3D Biomedical Segmentation

TL;DR

Abstract

nnActive: A Framework for Evaluation of Active Learning in 3D Biomedical Segmentation

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (25)