Table of Contents
Fetching ...

A Stochastic Algorithm for Searching Saddle Points with Convergence Guarantee

Baoming Shi, Lei Zhang, Qiang Du

TL;DR

This work develops and analyzes a stochastic saddle-point search method that circumvents exact gradient and Hessian evaluations by employing stochastic approximations of unstable directions. It establishes global almost-sure convergence for the variant with a known unstable space and convex–concave structure, and local almost-sure convergence (with an $O(1/n)$ rate) when the unstable directions are known only approximately via a stochastic eigenvector search. The authors prove that the stochastic eigenvector-search converges to the relevant eigenspace and show local high-probability convergence for the overall algorithm when unstable directions are estimated, with performance governed by the Hessian structure and gradient-noise level. Numerical experiments on a Müller–Brown potential, butterfly energy landscape, neural-network loss landscapes, and the Landau–de Gennes energy functional demonstrate practical effectiveness, including escaping from bad regions and achieving convergence in challenging high-dimensional or degenerate settings.

Abstract

Saddle points provide a hierarchical view of the energy landscape, revealing transition pathways and interconnected basins of attraction, and offering insight into the global structure, metastability, and possible collective mechanisms of the underlying system. In this work, we propose a stochastic saddle-search algorithm to circumvent exact derivative and Hessian evaluations that have been used in implementing traditional and deterministic saddle dynamics. At each iteration, the algorithm uses a stochastic eigenvector-search method, based on a stochastic Hessian, to approximate the unstable directions, followed by a stochastic gradient update with reflections in the approximate unstable direction to advance toward the saddle point. We carry out rigorous numerical analysis to establish the almost sure convergence for the stochastic eigenvector search and local almost sure convergence with an $O(1/n)$ rate for the saddle search, and present a theoretical guarantee to ensure the high-probability identification of the saddle point when the initial point is sufficiently close. Numerical experiments, including the application to a neural network loss landscape and a Landau-de Gennes type model for nematic liquid crystal, demonstrate the practical applicability and the ability for escaping from "bad" areas of the algorithm.

A Stochastic Algorithm for Searching Saddle Points with Convergence Guarantee

TL;DR

This work develops and analyzes a stochastic saddle-point search method that circumvents exact gradient and Hessian evaluations by employing stochastic approximations of unstable directions. It establishes global almost-sure convergence for the variant with a known unstable space and convex–concave structure, and local almost-sure convergence (with an rate) when the unstable directions are known only approximately via a stochastic eigenvector search. The authors prove that the stochastic eigenvector-search converges to the relevant eigenspace and show local high-probability convergence for the overall algorithm when unstable directions are estimated, with performance governed by the Hessian structure and gradient-noise level. Numerical experiments on a Müller–Brown potential, butterfly energy landscape, neural-network loss landscapes, and the Landau–de Gennes energy functional demonstrate practical effectiveness, including escaping from bad regions and achieving convergence in challenging high-dimensional or degenerate settings.

Abstract

Saddle points provide a hierarchical view of the energy landscape, revealing transition pathways and interconnected basins of attraction, and offering insight into the global structure, metastability, and possible collective mechanisms of the underlying system. In this work, we propose a stochastic saddle-search algorithm to circumvent exact derivative and Hessian evaluations that have been used in implementing traditional and deterministic saddle dynamics. At each iteration, the algorithm uses a stochastic eigenvector-search method, based on a stochastic Hessian, to approximate the unstable directions, followed by a stochastic gradient update with reflections in the approximate unstable direction to advance toward the saddle point. We carry out rigorous numerical analysis to establish the almost sure convergence for the stochastic eigenvector search and local almost sure convergence with an rate for the saddle search, and present a theoretical guarantee to ensure the high-probability identification of the saddle point when the initial point is sufficiently close. Numerical experiments, including the application to a neural network loss landscape and a Landau-de Gennes type model for nematic liquid crystal, demonstrate the practical applicability and the ability for escaping from "bad" areas of the algorithm.
Paper Structure (14 sections, 17 theorems, 96 equations, 4 figures, 2 algorithms)

This paper contains 14 sections, 17 theorems, 96 equations, 4 figures, 2 algorithms.

Key Result

Proposition 3.1

\newlabelLyapunov0 Suppose convex-concave holds, the stationary point of eq: simple SD is Lyapunov stable. Moreover, with any initial condition, the solution $(\hat{\mathbf{x}_{\mathcal{V}}}(t),\hat{\mathbf{x}_{\mathcal{V}{\space\perp}}}(t))$ of eq: simple SD asymptotically converges to the saddle

Figures (4)

  • Figure 1: (a) Contour plot of MB potential and iteration points of one-time stochastic saddle-search algorithm with $\nabla f(\mathbf{x};\omega)=\nabla f(\mathbf{x})+100\xi, \xi\sim\mathcal{N}(\mathbf{0},\mathbf{I}_d)$. (b) Average error $\Vert \mathbf{x}(n)-\mathbf{x}^* \Vert^2$ over 100 runs of the stochastic saddle‑search algorithm with $\alpha(n)=0.01/(n+100)$ (black) and $\alpha(n)\equiv0.0001$ (blue).
  • Figure 2: Contour plot of the butterfly energy landscape, together with trajectories of the stochastic saddle‑search algorithm ($\nabla f(\mathbf{x};\omega)=\nabla f(\mathbf{x})+\xi, \xi\sim\mathcal{N}(\mathbf{0},\mathbf{I}_d)$) and the deterministic saddle dynamics with $\alpha(n)=0.5/n$. The black quivers indicate the direction of the modified gradient with reflection in the eigenvector associated with the smallest eigenvalue, i.e., $-(\mathbf{I}_2-2 \mathbf{v} \mathbf{v}^\top)\nabla f(\mathbf{x})$. The olive‑colored region roughly depicts the domain of attraction of the saddle point.
  • Figure 3: Plots of the squared gradient norm $\|\nabla f(\mathbf{W}(n))\|_2^2$ versus the iteration number $n$ with decay step size $\alpha(n)=-100/(n+10000)$, for the neural network loss function. (a) The dataset size is $N = 100$, with a mini-batch size $|I_n|=20$. (b) $N = 10000$, $|I_n|=1000$.
  • Figure 4: (a) Transition pathway between two dual stable states D$_1$ and D$_2$ passing through the index‑1 transition state BD. The color bar labels the order parameter $\sqrt{|\mathbf{Q}|^2/2}$ and the white lines represent the directors, i.e., the eigenvector corresponding to the largest eigenvalue of $\mathbf{Q}$. (b) The error $\lVert\nabla E(\mathbf{Q}(n))\rVert^2_2$ versus iteration count $n$ for the stochastic saddle‑search algorithm with step size $\alpha(n)=1/(n+10000)$ and an initial condition near the D$_1$ state.

Theorems & Definitions (47)

  • Proposition 3.1
  • Proof 1
  • Definition 3.2
  • Lemma 3.3
  • Proof 2
  • Proposition 3.4: Boundedness
  • Proof 3
  • Proposition 3.5
  • Proof 4
  • Remark 4.1
  • ...and 37 more