A Stochastic Algorithm for Searching Saddle Points with Convergence Guarantee
Baoming Shi, Lei Zhang, Qiang Du
TL;DR
This work develops and analyzes a stochastic saddle-point search method that circumvents exact gradient and Hessian evaluations by employing stochastic approximations of unstable directions. It establishes global almost-sure convergence for the variant with a known unstable space and convex–concave structure, and local almost-sure convergence (with an $O(1/n)$ rate) when the unstable directions are known only approximately via a stochastic eigenvector search. The authors prove that the stochastic eigenvector-search converges to the relevant eigenspace and show local high-probability convergence for the overall algorithm when unstable directions are estimated, with performance governed by the Hessian structure and gradient-noise level. Numerical experiments on a Müller–Brown potential, butterfly energy landscape, neural-network loss landscapes, and the Landau–de Gennes energy functional demonstrate practical effectiveness, including escaping from bad regions and achieving convergence in challenging high-dimensional or degenerate settings.
Abstract
Saddle points provide a hierarchical view of the energy landscape, revealing transition pathways and interconnected basins of attraction, and offering insight into the global structure, metastability, and possible collective mechanisms of the underlying system. In this work, we propose a stochastic saddle-search algorithm to circumvent exact derivative and Hessian evaluations that have been used in implementing traditional and deterministic saddle dynamics. At each iteration, the algorithm uses a stochastic eigenvector-search method, based on a stochastic Hessian, to approximate the unstable directions, followed by a stochastic gradient update with reflections in the approximate unstable direction to advance toward the saddle point. We carry out rigorous numerical analysis to establish the almost sure convergence for the stochastic eigenvector search and local almost sure convergence with an $O(1/n)$ rate for the saddle search, and present a theoretical guarantee to ensure the high-probability identification of the saddle point when the initial point is sufficiently close. Numerical experiments, including the application to a neural network loss landscape and a Landau-de Gennes type model for nematic liquid crystal, demonstrate the practical applicability and the ability for escaping from "bad" areas of the algorithm.
