Bregman Stochastic Proximal Point Algorithm with Variance Reduction
Cheik Traoré, Peter Ochs
TL;DR
The paper develops a unified variance-reduction framework for the Bregman stochastic proximal point algorithm (BSPPA) to address convergence slowdowns caused by stochasticity and non-Euclidean geometry. By introducing generic variance-reduction terms $e_k$ and deriving a robust convergence theory, it yields BSPPA variants that mimic SAGA and SVRG (BSAPA, BSVRP, and BLSVRP) and achieve improved sublinear and linear rates under relative smoothness and relative strong convexity, without requiring vanishing stepsizes. The authors provide detailed instantiations, theoretical guarantees, and experimental validation on Poisson KL inverse problems and tomographic reconstruction, demonstrating enhanced stability and faster convergence compared to vanilla BSPPA and SPPA. The work also unifies the analysis with Bregman SGD, enabling a broad, non-Euclidean perspective on variance-reduced stochastic optimization for constrained or non-Lipschitz settings, with future extensions to nonsmooth and inexact proximal mappings.
Abstract
Stochastic algorithms, especially stochastic gradient descent (SGD), have proven to be the go-to methods in data science and machine learning. In recent years, the stochastic proximal point algorithm (SPPA) emerged, and it was shown to be more robust than SGD with respect to stepsize settings. However, SPPA still suffers from a decreased convergence rate due to the need for vanishing stepsizes, which is resolved by using variance reduction methods. In the deterministic setting, there are many problems that can be solved more efficiently when viewing them in a non-Euclidean geometry using Bregman distances. This paper combines these two worlds and proposes variance reduction techniques for the Bregman stochastic proximal point algorithm (BSPPA). As special cases, we obtain SAGA- and SVRG-like variance reduction techniques for BSPPA. Our theoretical and numerical results demonstrate improved stability and convergence rates compared to the vanilla BSPPA with constant and vanishing stepsizes, respectively. Our analysis, also, allow to recover the same variance reduction techniques for Bregman SGD in a unified way.
