Table of Contents
Fetching ...

On the convergence of stochastic variance reduced gradient for linear inverse problems

Bangti Jin, Zehui Zhou

TL;DR

This paper analyzes stochastic variance reduced gradient methods for linear inverse problems in Hilbert spaces, focusing on SVRG and a regularized variant $r$SVRG that incorporates an approximate operator $A$ to inject a learned prior. Under a fixed step-size and a source condition on the initial error, the authors prove that $r$SVRG achieves the optimal noise-dependent convergence rate without early stopping, while SVRG attains the same rate with a suitable a priori stopping for nonsmooth solutions. The results quantify when a learned or truncated operator regularizes the problem and improves convergence, and are supported by numerical tests on three discretized ill-posed problems showing the practical advantages of $r$SVRG. The work also outlines conditions under which the convergence rates hold in expectation and uniformly, and discusses relaxations relative to prior analyses.

Abstract

Stochastic variance reduced gradient (SVRG) is an accelerated version of stochastic gradient descent based on variance reduction, and is promising for solving large-scale inverse problems. In this work, we analyze SVRG and a regularized version that incorporates a priori knowledge of the problem, for solving linear inverse problems in Hilbert spaces. We prove that, with suitable constant step size schedules and regularity conditions, the regularized SVRG can achieve optimal convergence rates in terms of the noise level without any early stopping rules, and standard SVRG is also optimal for problems with nonsmooth solutions under a priori stopping rules. The analysis is based on an explicit error recursion and suitable prior estimates on the inner loop updates with respect to the anchor point. Numerical experiments are provided to complement the theoretical analysis.

On the convergence of stochastic variance reduced gradient for linear inverse problems

TL;DR

This paper analyzes stochastic variance reduced gradient methods for linear inverse problems in Hilbert spaces, focusing on SVRG and a regularized variant SVRG that incorporates an approximate operator to inject a learned prior. Under a fixed step-size and a source condition on the initial error, the authors prove that SVRG achieves the optimal noise-dependent convergence rate without early stopping, while SVRG attains the same rate with a suitable a priori stopping for nonsmooth solutions. The results quantify when a learned or truncated operator regularizes the problem and improves convergence, and are supported by numerical tests on three discretized ill-posed problems showing the practical advantages of SVRG. The work also outlines conditions under which the convergence rates hold in expectation and uniformly, and discusses relaxations relative to prior analyses.

Abstract

Stochastic variance reduced gradient (SVRG) is an accelerated version of stochastic gradient descent based on variance reduction, and is promising for solving large-scale inverse problems. In this work, we analyze SVRG and a regularized version that incorporates a priori knowledge of the problem, for solving linear inverse problems in Hilbert spaces. We prove that, with suitable constant step size schedules and regularity conditions, the regularized SVRG can achieve optimal convergence rates in terms of the noise level without any early stopping rules, and standard SVRG is also optimal for problems with nonsmooth solutions under a priori stopping rules. The analysis is based on an explicit error recursion and suitable prior estimates on the inner loop updates with respect to the anchor point. Numerical experiments are provided to complement the theoretical analysis.
Paper Structure (8 sections, 12 theorems, 94 equations, 1 figure, 3 tables, 2 algorithms)

This paper contains 8 sections, 12 theorems, 94 equations, 1 figure, 3 tables, 2 algorithms.

Key Result

Theorem 2.1

Let Assumption ass hold with $b=(1+2\nu)^{-1}$. Then there exists some $c^*$ independent of $k$, $n$ or $\delta$ such that, for any $k\geq 0$,

Figures (1)

  • Figure 4.1: The convergence of the relative error $e={\mathbb{E}[\| x_{k}^\delta-x_\dag\|^2]^\frac{1}{2}}/{\|x_\dag\|}$ versus the iteration number $k$ for phillips, gravity and shaw. The rows from top to bottom are for $\epsilon=$1e-3, $\epsilon=$5e-3, $\epsilon=$1e-2 and $\epsilon=$5e-2, respectively. The intersection of the gray dashed lines represents the stopping point, determined by the discrepancy principle, for LM along the iteration trajectory.

Theorems & Definitions (26)

  • Theorem 2.1
  • Corollary 2.1
  • Remark 2.1
  • Corollary 2.2
  • Corollary 2.3
  • Lemma 3.1
  • proof
  • Lemma 3.2
  • proof
  • Theorem 3.1
  • ...and 16 more