Table of Contents
Fetching ...

A Structured Review and Quantitative Profiling of Public Brain MRI Datasets for Foundation Model Development

Minh Sao Khue Luu, Margaret V. Benedichuk, Ekaterina I. Roppert, Roman M. Kenzhin, Bair N. Tuchinov

TL;DR

We survey 54 public brain MRI datasets to characterize dataset- and image-level variability relevant to large-scale representation learning from MRI. The study standardizes modalities and cohorts, catalogs metadata, and analyzes disease coverage, dataset scale, modality composition, voxel geometry, and intensity distributions. It assesses preprocessing effects on harmonization and demonstrates residual covariate shift in feature space after standardized processing, arguing that preprocessing alone cannot erase inter-dataset bias. The findings highlight the need for preprocessing-aware and domain-adaptive strategies to build robust, generalizable brain MRI representations from heterogeneous public resources.

Abstract

The development of foundation models for brain MRI depends critically on the scale, diversity, and consistency of available data, yet systematic assessments of these factors remain scarce. In this study, we analyze 54 publicly accessible brain MRI datasets encompassing over 538,031 to provide a structured, multi-level overview tailored to foundation model development. At the dataset level, we characterize modality composition, disease coverage, and dataset scale, revealing strong imbalances between large healthy cohorts and smaller clinical populations. At the image level, we quantify voxel spacing, orientation, and intensity distributions across 15 representative datasets, demonstrating substantial heterogeneity that can influence representation learning. We then perform a quantitative evaluation of preprocessing variability, examining how intensity normalization, bias field correction, skull stripping, spatial registration, and interpolation alter voxel statistics and geometry. While these steps improve within-dataset consistency, residual differences persist between datasets. Finally, feature-space case study using a 3D DenseNet121 shows measurable residual covariate shift after standardized preprocessing, confirming that harmonization alone cannot eliminate inter-dataset bias. Together, these analyses provide a unified characterization of variability in public brain MRI resources and emphasize the need for preprocessing-aware and domain-adaptive strategies in the design of generalizable brain MRI foundation models.

A Structured Review and Quantitative Profiling of Public Brain MRI Datasets for Foundation Model Development

TL;DR

We survey 54 public brain MRI datasets to characterize dataset- and image-level variability relevant to large-scale representation learning from MRI. The study standardizes modalities and cohorts, catalogs metadata, and analyzes disease coverage, dataset scale, modality composition, voxel geometry, and intensity distributions. It assesses preprocessing effects on harmonization and demonstrates residual covariate shift in feature space after standardized processing, arguing that preprocessing alone cannot erase inter-dataset bias. The findings highlight the need for preprocessing-aware and domain-adaptive strategies to build robust, generalizable brain MRI representations from heterogeneous public resources.

Abstract

The development of foundation models for brain MRI depends critically on the scale, diversity, and consistency of available data, yet systematic assessments of these factors remain scarce. In this study, we analyze 54 publicly accessible brain MRI datasets encompassing over 538,031 to provide a structured, multi-level overview tailored to foundation model development. At the dataset level, we characterize modality composition, disease coverage, and dataset scale, revealing strong imbalances between large healthy cohorts and smaller clinical populations. At the image level, we quantify voxel spacing, orientation, and intensity distributions across 15 representative datasets, demonstrating substantial heterogeneity that can influence representation learning. We then perform a quantitative evaluation of preprocessing variability, examining how intensity normalization, bias field correction, skull stripping, spatial registration, and interpolation alter voxel statistics and geometry. While these steps improve within-dataset consistency, residual differences persist between datasets. Finally, feature-space case study using a 3D DenseNet121 shows measurable residual covariate shift after standardized preprocessing, confirming that harmonization alone cannot eliminate inter-dataset bias. Together, these analyses provide a unified characterization of variability in public brain MRI resources and emphasize the need for preprocessing-aware and domain-adaptive strategies in the design of generalizable brain MRI foundation models.
Paper Structure (23 sections, 12 figures, 7 tables)

This paper contains 23 sections, 12 figures, 7 tables.

Figures (12)

  • Figure 1: Distribution of subjects by disease category after removing the undefined "Multiple Diseases" group. The x-axis uses a logarithmic scale to enable visualization across several orders of magnitude, from hundreds to tens of thousands of subjects.
  • Figure 2: Distribution of dataset sizes on a logarithmic scale. The figure highlights the dominance of one extremely large population dataset and numerous smaller, clinically focused cohorts. Logarithmic scaling compresses large numerical differences to emphasize structural imbalance across dataset scales.
  • Figure 3: Heatmap of modality co-occurrence across structural MRI datasets. High-intensity cells indicate frequent pairing between modalities, particularly T1–FLAIR, T1–T2, and T1–T1c. These patterns reveal partial but consistent overlap that supports unified representation learning across multi-dataset collections.
  • Figure 4: Voxel spacing distribution (in mm) along the $x$, $y$, and $z$ axes for 14 curated datasets. Each point represents one scan, and each color corresponds to a dataset. Compact clusters indicate consistent acquisition protocols, while spread-out points show variation in resolution and anisotropy.
  • Figure 5: Median voxel intensity per image across datasets. Each dot represents one 3D MRI volume.
  • ...and 7 more figures