Table of Contents
Fetching ...

Physics-guided Emulators Reveal Resilience and Fragility under Operational Latencies and Outages

Sarth Dubey, Subimal Ghosh, Udit Bhatia

TL;DR

This work tackles the challenge of reliable hydrologic forecasting under operational data constraints by developing a physics-guided emulator of the GloFAS core. The authors implement a 365-day lag with a 10-day lead encoder–decoder LSTM, trained with a soft water-balance constraint to preserve physical coherence, and evaluate five latency-aware architectures across data-rich and data-scarce basins. They demonstrate that predictive skill degrades gradually, not catastrophically, as data latency increases, and that short-range forecasts can recover some lost performance; cross-domain transfer reveals robust generalization within data-rich regimes but limits under heavy regulation and data scarcity. Collectively, the results establish operational robustness as a measurable, designable property of hydrologic machine learning and offer a framework for evaluating real-time forecasting systems under imperfect data streams.

Abstract

Reliable hydrologic and flood forecasting requires models that remain stable when input data are delayed, missing, or inconsistent. However, most advances in rainfall-runoff prediction have been evaluated under ideal data conditions, emphasizing accuracy rather than operational resilience. Here, we develop an operationally ready emulator of the Global Flood Awareness System (GloFAS) that couples long- and short-term memory networks with a relaxed water-balance constraint to preserve physical coherence. Five architectures span a continuum of information availability: from complete historical and forecast forcings to scenarios with data latency and outages, allowing systematic evaluation of robustness. Trained in minimally managed catchments across the United States and tested in more than 5,000 basins, including heavily regulated rivers in India, the emulator reproduces the hydrological core of GloFAS and degrades smoothly as information quality declines. Transfer across contrasting hydroclimatic and management regimes yields reduced yet physically consistent performance, defining the limits of generalization under data scarcity and human influence. The framework establishes operational robustness as a measurable property of hydrological machine learning and advances the design of reliable real-time forecasting systems.

Physics-guided Emulators Reveal Resilience and Fragility under Operational Latencies and Outages

TL;DR

This work tackles the challenge of reliable hydrologic forecasting under operational data constraints by developing a physics-guided emulator of the GloFAS core. The authors implement a 365-day lag with a 10-day lead encoder–decoder LSTM, trained with a soft water-balance constraint to preserve physical coherence, and evaluate five latency-aware architectures across data-rich and data-scarce basins. They demonstrate that predictive skill degrades gradually, not catastrophically, as data latency increases, and that short-range forecasts can recover some lost performance; cross-domain transfer reveals robust generalization within data-rich regimes but limits under heavy regulation and data scarcity. Collectively, the results establish operational robustness as a measurable, designable property of hydrologic machine learning and offer a framework for evaluating real-time forecasting systems under imperfect data streams.

Abstract

Reliable hydrologic and flood forecasting requires models that remain stable when input data are delayed, missing, or inconsistent. However, most advances in rainfall-runoff prediction have been evaluated under ideal data conditions, emphasizing accuracy rather than operational resilience. Here, we develop an operationally ready emulator of the Global Flood Awareness System (GloFAS) that couples long- and short-term memory networks with a relaxed water-balance constraint to preserve physical coherence. Five architectures span a continuum of information availability: from complete historical and forecast forcings to scenarios with data latency and outages, allowing systematic evaluation of robustness. Trained in minimally managed catchments across the United States and tested in more than 5,000 basins, including heavily regulated rivers in India, the emulator reproduces the hydrological core of GloFAS and degrades smoothly as information quality declines. Transfer across contrasting hydroclimatic and management regimes yields reduced yet physically consistent performance, defining the limits of generalization under data scarcity and human influence. The framework establishes operational robustness as a measurable property of hydrological machine learning and advances the design of reliable real-time forecasting systems.
Paper Structure (38 sections, 11 equations, 15 figures, 1 table)

This paper contains 38 sections, 11 equations, 15 figures, 1 table.

Figures (15)

  • Figure 1: Architecture and experimental design of the data-latency-aware emulator. The schematic integrates the encoder–decoder structure with operational availability masks and the domain hierarchy used in this study. The upper maps delineate the three domains—CAMELS-US (minimally managed, 395 catchments), HYSETS (managed but data-rich, 5,149 catchments) and CAMELS-IND (heavily managed and data-scarce, 191 catchments)—that form the source and target settings for transfer experiments. The middle panels illustrate the long-short-term-memory (LSTM) encoder assimilating 365 days of ERA5, GPM, and static attributes and the decoder projecting 10-day discharge and soil-wetness forecasts. The lower block defines the operational availability mask that governs which meteorological forcings reach the decoder in historical or operational mode, including cases with delayed or missing forecasts. Together these elements constitute a physics-guided, latency-aware surrogate of the GloFAS hydrological core designed to evaluate robustness under realistic data constraints
  • Figure 2: Regional to generalizable surrogacy within data-rich domains. Cross-HUC performance matrices (top) show that models trained and tested within the same region reproduce discharge behaviour faithfully (median NSE $>$ 0.6; F1 $>$ 0.7), while moderate off-diagonal skill (NSE $\sim$ 0.3–0.5) indicates transferable hydrologic structure. Cumulative distributions (middle) reveal that a single continental model trained across all U.S. basins performs as well as, or slightly better than, region-specific models, confirming that large-sample training enhances stability. Skill declines systematically with increasing flow intermittency (bottom), from NSE $\sim$ 0.8 in perennial rivers to < 0 in highly intermittent basins, exposing the weak rainfall–runoff coupling that limits model generalization. The figure establishes that continental-scale learning yields a stable surrogate for GloFAS while revealing the hydroclimatic regimes that define its limits
  • Figure 3: Surrogacy and zero-shot transfer under full data availability. Spatial maps of NSE and F1 (a–f) show that the emulator reproduces GloFAS discharge skill across the United States (median NSE $\sim$ 0.7; F1 $\sim$ 0.8), maintains coherent performance in HYSETS ($\sim$ 0.55 and 0.7) and retains structured, though weaker, skill in CAMELS-IND ($\sim$ 0.4 and 0.6). Lead-time profiles (g–l) remain essentially flat over ten days ($\Delta$ NSE $<$ 0.03; $\Delta$ F1 $<$ 0.02), demonstrating that the model evolves a continuous hydrologic state rather than compounding forecast errors. These results confirm that the emulator captures the hydrological logic of GloFAS and generalizes across contrasting hydroclimates, defining the upper bound of achievable emulation under complete data availability
  • Figure 4: Operational performance under data latency and outages. Probability and cumulative distributions of F1 (top six panels) compare model skill when the decoder ingests ECMWF-HRES forecasts (Using HRES, a–c) and when forecasts are withheld (No Meteorological Forecasts, d–f). Median F1 drops only slightly from 0.81 to 0.78 in CAMELS-US, 0.74 to 0.70 in HYSETS, and 0.62 to 0.58 in CAMELS-IND; corresponding NSE reductions are similarly small ($\sim$ 0.03). The schematic (g) depicts how the encoder assimilates 365 days of hydro-meteorological history while the decoder adapts to the presence or absence of forecasts. The model therefore remains stable and physically coherent under degraded inputs, providing a quantitative measure of operational robustness and a benchmark for testing alternative architectures
  • Figure 5: Quantification of performance degradation across architectures and forecast horizons. Distributions of Nash–Sutcliffe efficiency and F1 score (top two rows) show a smooth decline from the full-data configuration (H1) to the most degraded (H4) and partial recovery when short-range forecasts are added (H5), with the steepest losses in data-scarce, heavily managed basins. Lead-time profiles (bottom two rows) remain stable for configurations with consistent meteorological input but deteriorate beyond day 5 when forecasts are absent, establishing operational robustness as a measurable property of the emulator.
  • ...and 10 more figures