Table of Contents
Fetching ...

Extreme Event Aware ($η$-) Learning

Kai Chang, Themistoklis P. Sapsis

TL;DR

Extreme events are rare yet impactful, and standard data-driven methods struggle when extremes are underrepresented. The authors formulate $η$-learning, an optimization framework that augments supervised losses with a 1-Wasserstein regularization term against a prescribed tail distribution $ν_0$, enabling the generation of unprecedented extremes consistent with physics. Grounded in optimal transport theory, the approach provides data-consistency limits and shows near-optimal tail-statistics performance, addressing the core limitations of data-scarce regime. Through toy experiments and real-world precipitation downscaling, $η$-learning demonstrates improved tail fidelity, offering a model-agnostic post-processing tool for enhancing extreme-event statistics in diverse domains.

Abstract

Quantifying and predicting rare and extreme events persists as a crucial yet challenging task in understanding complex dynamical systems. Many practical challenges arise from the infrequency and severity of these events, including the considerable variance of simple sampling methods and the substantial computational cost of high-fidelity numerical simulations. Numerous data-driven methods have recently been developed to tackle these challenges. However, a typical assumption for the success of these methods is the occurrence of multiple extreme events, either within the training dataset or during the sampling process. This leads to accurate models in regions of quiescent events but with high epistemic uncertainty in regions associated with extremes. To overcome this limitation, we introduce Extreme Event Aware (e2a or eta) or $η$-learning which does not assume the existence of extreme events in the available data. $η$-learning reduces the uncertainty even in `uncharted' extreme event regions, by enforcing the extreme event statistics of an observable indicative of extremeness during training, which can be available through qualitative arguments or estimated with unlabeled data. This type of statistical regularization results in models that fit the observed data, while enforcing consistency with the prescribed observable statistics, enabling the generation of unprecedented extreme events even when the training data lack extremes therein. Theoretical results based on optimal transport offer a rigorous justification and highlight the optimality of the introduced method. Additionally, extensive numerical experiments illustrate the favorable properties of the $η$-learning framework on several prototype problems and real-world precipitation downscaling problems.

Extreme Event Aware ($η$-) Learning

TL;DR

Extreme events are rare yet impactful, and standard data-driven methods struggle when extremes are underrepresented. The authors formulate -learning, an optimization framework that augments supervised losses with a 1-Wasserstein regularization term against a prescribed tail distribution , enabling the generation of unprecedented extremes consistent with physics. Grounded in optimal transport theory, the approach provides data-consistency limits and shows near-optimal tail-statistics performance, addressing the core limitations of data-scarce regime. Through toy experiments and real-world precipitation downscaling, -learning demonstrates improved tail fidelity, offering a model-agnostic post-processing tool for enhancing extreme-event statistics in diverse domains.

Abstract

Quantifying and predicting rare and extreme events persists as a crucial yet challenging task in understanding complex dynamical systems. Many practical challenges arise from the infrequency and severity of these events, including the considerable variance of simple sampling methods and the substantial computational cost of high-fidelity numerical simulations. Numerous data-driven methods have recently been developed to tackle these challenges. However, a typical assumption for the success of these methods is the occurrence of multiple extreme events, either within the training dataset or during the sampling process. This leads to accurate models in regions of quiescent events but with high epistemic uncertainty in regions associated with extremes. To overcome this limitation, we introduce Extreme Event Aware (e2a or eta) or -learning which does not assume the existence of extreme events in the available data. -learning reduces the uncertainty even in `uncharted' extreme event regions, by enforcing the extreme event statistics of an observable indicative of extremeness during training, which can be available through qualitative arguments or estimated with unlabeled data. This type of statistical regularization results in models that fit the observed data, while enforcing consistency with the prescribed observable statistics, enabling the generation of unprecedented extreme events even when the training data lack extremes therein. Theoretical results based on optimal transport offer a rigorous justification and highlight the optimality of the introduced method. Additionally, extensive numerical experiments illustrate the favorable properties of the -learning framework on several prototype problems and real-world precipitation downscaling problems.
Paper Structure (39 sections, 18 theorems, 86 equations, 18 figures, 2 algorithms)

This paper contains 39 sections, 18 theorems, 86 equations, 18 figures, 2 algorithms.

Key Result

Theorem 4

Suppose both $y$ and $y_{\xi}$ are in $L^1(\mu)$ and are $\mathcal{B}(\mathcal{X})$-measurable with respect to $\mathcal{B}(\mathcal{Y})$. Then for a data-consistent estimator $y_{\xi}$ in the sense of eq:data-consistent, when $n\leq {\log p \over \log{(1-\delta)}}$, with probability at least $p$, w where $\tilde{C}$ is the same as that in eq:data-consistent.

Figures (18)

  • Figure 1: The contour plot comparison of different estimators against the ground truth evaluated on $[-6,6] \times [-6,6]$ for the 2D-to-1D toy example. Left: ground truth with the crosses marking the training data that does not contain the extreme of interest. Middle: the MSE estimator trained with the selected training data. Right: the $\eta$-estimator trained with the same training data along with a reference distribution.
  • Figure 2: The contour plot comparison of different estimators against the ground truth evaluated on $[-6,6] \times [-6,6]$ for the 2D-to-1D toy example. Left column: ground truth with the crosses marking the training data that does not contain the extreme of interest. Middle column: the MSE estimator trained with the selected training data. Right column: the $\eta$-estimator trained with the same training data along with a reference distribution.
  • Figure 3: Visualization of sample precipitation snapshots. Each column corresponds to a particular day in the test set. Each row corresponds to a particular processing method. Top row: downsampled precipitation fields. Second to the top row: ground truth HR fields. Third to the top row: HR fields downscaled from the first row by the MSE map. Bottom row: HR fields downscaled from the first row by the $\eta$-map.
  • Figure 4: The observable PDF comparison of HR daily maximal daily peak precipitation fields under different processing methods. The solid blue curve: the 25-year ground truth high-resolution (HR) data. The dotted orange curve with rhombus markers: the 0.5-year HR training data. The dotted green curve with triangle markers: the HR samples super-resoluted from the 25-year ground truth LR samples through the MSE map. The dotted red curve with star markers: the HR samples super-resoluted from the 25-year ground truth LR samples through the $\eta$-map.
  • Figure 5: The observable PDF comparison of low-resolution (LR) and high-resolution (HR) maximal daily peak precipitation fields under different processing methods. The solid blue curve: the 25-year ground truth HR data. The dotted purple curve with square markers: 25-year ground truth LR data. The dotted green curve with triangle markers: the LR samples drawn from a Flow Matching (FM) model trained with 25-year LR data. The dotted orange curve with rhombus markers: the 0.5-year HR data used to train the HR FM model. The dotted olive curve with circle markers: the samples drawn from the mentioned HR FM model. The dotted red curve with star markers: the HR samples generated from the $\eta$-map by mapping the samples drawn from the LR FM model through the $\eta$-map.
  • ...and 13 more figures

Theorems & Definitions (32)

  • Definition 1: Empirical Risk Minimization (Chapter 4 of bach2024learning)
  • Definition 2: 1-Wasserstein Distance sot
  • Definition 3: Data-Consistent Estimator
  • Theorem 4: Lower Bound
  • Theorem 6
  • Theorem 7: Optimality of $W_1$ in Terms of Relative Approximation Errors
  • Theorem 9: Optimality of $W_1$ in Terms of Exact Approximation Error
  • Theorem 10: Kantorovich Duality of $W_1$ (Theorem 1.21 in sot)
  • Lemma 11
  • proof
  • ...and 22 more