Improved Central Limit Theorem and Bootstrap Approximations for Linear Stochastic Approximation
Bogdan Butyrin, Eric Moulines, Alexey Naumov, Sergey Samsonov, Qi-Man Shao, Zhuo-Song Zhang
TL;DR
This work advances finite-sample understanding of linear stochastic approximation by establishing a Berry–Esseen-type bound for the multivariate normal approximation of Polyak–Ruppert averaged iterates with decreasing step sizes, achieving a rate of $n^{-1/3}$ in convex distance when the step sizes decay as $k^{-rac{2}{3}}$. It introduces a non-asymptotic multiplier bootstrap procedure that mirrors the distribution of the rescaled averaged estimator, with an approximation rate up to $n^{-1/2}$ (up to logarithmic factors) under suitable growth of the step-size and sample size. The analysis improves prior results (e.g., Samsonov et al., 2024) by leveraging a Gaussian approximation based on $ abla_n$ rather than the limiting covariance, enabling sharper finite-sample guarantees and practical confidence-set construction. The results also connect to TD learning and broader SGD-type algorithms, offering a principled non-asymptotic framework for inference in linear stochastic approximation with explicit dependencies on dimension and step-size parameters. The work leaves open extensions to Markov noise settings and tight lower bounds in certain regimes, pointing to future research directions in non-iid data scenarios and optimality questions for step-size schedules.
Abstract
In this paper, we refine the Berry-Esseen bounds for the multivariate normal approximation of Polyak-Ruppert averaged iterates arising from the linear stochastic approximation (LSA) algorithm with decreasing step size. We consider the normal approximation by the Gaussian distribution with covariance matrix predicted by the Polyak-Juditsky central limit theorem and establish the rate up to order $n^{-1/3}$ in convex distance, where $n$ is the number of samples used in the algorithm. We also prove a non-asymptotic validity of the multiplier bootstrap procedure for approximating the distribution of the rescaled error of the averaged LSA estimator. We establish approximation rates of order up to $1/\sqrt{n}$ for the latter distribution, which significantly improves upon the previous results obtained by Samsonov et al. (2024).
