SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation

Kai-Hendrik Cohrs; Zuzanna Osika; Maria Gonzalez-Calabuig; Vishal Nedungadi; Ruben Cartuyvels; Steffen Knoblauch; Joppe Massant; Shruti Nath; Patrick Ebel; Vasileios Sitokonstantinou

SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation

Kai-Hendrik Cohrs, Zuzanna Osika, Maria Gonzalez-Calabuig, Vishal Nedungadi, Ruben Cartuyvels, Steffen Knoblauch, Joppe Massant, Shruti Nath, Patrick Ebel, Vasileios Sitokonstantinou

TL;DR

SHRUG-FM addresses the reliability gap of geospatial foundation models under distribution shifts by integrating three complementary signals: input-space OOD, embedding-space OOD, and task-specific predictive uncertainty. Using burn scar segmentation with SSL4EO-S12 encodings and HydroATLAS context, it demonstrates that OOD scores correlate with degraded performance and that uncertainty-based flags can safely discard many unreliable predictions. The approach employs frozen foundation encoders plus a downstream decoder and ensembles (Deep Ensembles and MC Dropout) to quantify uncertainty at pixel and image levels, with thorough metric definitions for calibration and reliability. By linking failures to geospatial attributes and providing a dashboard for interpretable reliability assessment, SHRUG-FM offers a practical pathway toward safer deployment of GFMs in climate-sensitive applications and informs data-pretraining strategies to reduce future gaps.

Abstract

Geospatial foundation models for Earth observation often fail to perform reliably in environments underrepresented during pretraining. We introduce SHRUG-FM, a framework for reliability-aware prediction that integrates three complementary signals: out-of-distribution (OOD) detection in the input space, OOD detection in the embedding space and task-specific predictive uncertainty. Applied to burn scar segmentation, SHRUG-FM shows that OOD scores correlate with lower performance in specific environmental conditions, while uncertainty-based flags help discard many poorly performing predictions. Linking these flags to land cover attributes from HydroATLAS shows that failures are not random but concentrated in certain geographies, such as low-elevation zones and large river areas, likely due to underrepresentation in pretraining data. SHRUG-FM provides a pathway toward safer and more interpretable deployment of GFMs in climate-sensitive applications, helping bridge the gap between benchmark performance and real-world reliability.

SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation

TL;DR

Abstract

SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (5)