Table of Contents
Fetching ...

Hyperparameter Selection via Early Stopping for Bayesian Semilinear PDEs

Maia Tienstra, Gottfried Hastermann

TL;DR

This paper addresses Bayesian inverse problems for semilinear PDEs by translating the nonlinear estimation task into a linear surrogate through a linearisation framework. By tuning the Gaussian prior's scale via an early-stopping discrepancy principle in the linearised setting, it achieves adaptive posterior contraction rates and frequentist coverage, which can be transferred back to the original nonlinear parameter through a Lipschitz solution map. The authors provide general theory and then instantiate it for the time-independent Schrödinger equation, both theoretically and numerically, using Ensemble Kalman Bucy Inversion to implement the data-driven prior tuning. The work thus offers a data-driven, computationally efficient method for near-optimal prior selection in nonlinear Bayesian PDE inverse problems, with practical validation on a classical benchmark problem.

Abstract

We study non-linear Bayesian inverse problems arising from semilinear partial differential equations (PDEs) that can be transformed into linear Bayesian inverse problems. We are then able to extend the early stopping for Ensemble Kalman-Bucy Filter (EnKBF) to these types of linearisable nonlinear problems as a way to tune the prior distribution. Using the linearisation method introduced in \cite{koers2024}, we transform the non-linear problem into a linear one, apply early stopping based on the discrepancy principle, and then pull back the resulting posterior to the posterior for the original parameter of interest. Following \cite{tienstra2025}, we show that this approach yields adaptive posterior contraction rates and frequentist coverage guarantees, under mild conditions on the prior covariance operator. From this, it immediately follows that Tikhonov regularisation coupled with the discrepancy principle contracts at the same rate. The proposed method thus provides a data-driven way to tune Gaussian priors via early stopping, which is both computationally efficient and statistically near optimal for nonlinear problems. Lastly, we demonstrate our results theoretically and numerically for the classical benchmark problem, the time-independent Schrödinger equation.

Hyperparameter Selection via Early Stopping for Bayesian Semilinear PDEs

TL;DR

This paper addresses Bayesian inverse problems for semilinear PDEs by translating the nonlinear estimation task into a linear surrogate through a linearisation framework. By tuning the Gaussian prior's scale via an early-stopping discrepancy principle in the linearised setting, it achieves adaptive posterior contraction rates and frequentist coverage, which can be transferred back to the original nonlinear parameter through a Lipschitz solution map. The authors provide general theory and then instantiate it for the time-independent Schrödinger equation, both theoretically and numerically, using Ensemble Kalman Bucy Inversion to implement the data-driven prior tuning. The work thus offers a data-driven, computationally efficient method for near-optimal prior selection in nonlinear Bayesian PDE inverse problems, with practical validation on a classical benchmark problem.

Abstract

We study non-linear Bayesian inverse problems arising from semilinear partial differential equations (PDEs) that can be transformed into linear Bayesian inverse problems. We are then able to extend the early stopping for Ensemble Kalman-Bucy Filter (EnKBF) to these types of linearisable nonlinear problems as a way to tune the prior distribution. Using the linearisation method introduced in \cite{koers2024}, we transform the non-linear problem into a linear one, apply early stopping based on the discrepancy principle, and then pull back the resulting posterior to the posterior for the original parameter of interest. Following \cite{tienstra2025}, we show that this approach yields adaptive posterior contraction rates and frequentist coverage guarantees, under mild conditions on the prior covariance operator. From this, it immediately follows that Tikhonov regularisation coupled with the discrepancy principle contracts at the same rate. The proposed method thus provides a data-driven way to tune Gaussian priors via early stopping, which is both computationally efficient and statistically near optimal for nonlinear problems. Lastly, we demonstrate our results theoretically and numerically for the classical benchmark problem, the time-independent Schrödinger equation.
Paper Structure (20 sections, 13 theorems, 116 equations, 1 figure, 1 algorithm)

This paper contains 20 sections, 13 theorems, 116 equations, 1 figure, 1 algorithm.

Key Result

Lemma 2.1

Let $c$ be continuously Fréchet differentiable on $B^{\mathcal{F}\times U}(f_0,u_{f_0}) \subseteq \mathcal{F}\times U$. Additionally, assume $D_2 c$ to be invertible and have a bounded inverse. Furthermore assume $Dc$ and $D_2^{-1}c$ to be a bounded linear operator on $\overline{B}^{\mathcal{F}\time and for every $v_f \in B^V(v_{f_0})$ and $f \in B^{\mathcal{F}}(f_0)$.

Figures (1)

  • Figure 1: Here we plot the results of running EKI with early stopping, \ref{['alg:deter_enkf']}, on the Schrödinger problem. From left to right, the noise level decreases. On the top is the estimation for $v_{0,i}$, the coefficients of $v_0$. We plot only the first $10$ coeffients as the remaining are essentially zero, and this zoomed-in perspective shows how uncertainty in the coefficients propagates to the uncertainty in $f_0$. On the bottom are the resulting transformed estimates for $f_0$ in function space over a grid of $100$ points, and $\kappa$ is chosen to be $D(n)*n$ where $D(n)=100$. For all noise levels, we fix the grid and only scale the variance of the noise in the linear observations. The red solid line is the ensemble mean, the black dashed line is the ground truth, the blue thinner dashed lines are the ensemble particles, and the blue filled region is the $95\%$ credible region computed by taking the $95\%$ quantiles of the ensemble.

Theorems & Definitions (37)

  • Example 1.1: Stationary Schrödinger Equation
  • Remark 2.1
  • Lemma 2.1
  • proof
  • proof
  • Example 2.1: Stationary Schrödinger Equation
  • Example 2.2: Darcy flow
  • Remark 2.2
  • Definition 2.1
  • Remark 2.3
  • ...and 27 more