Near-Optimal Tensor PCA via Normalized Stochastic Gradient Ascent with Overparameterization
Shihong Ding, Yihong Gu, Yuanshi Liu, Cong Fang
TL;DR
This work tackles tensor PCA in the spiked tensor model by proposing a normalized stochastic gradient ascent method with overparameterization (NSGA). It demonstrates that, without global or spectral initialization, the planted vector $v_*$ can be recovered with constant probability whenever $N\lambda^2\ge\widetilde{Ω}(d^{\lceil k/2\rceil})$ for the odd/even cases, achieving near-optimal sample complexity akin to $d^{k/2}$ thresholds known from SoS relaxations. The analysis reveals two-phase learning dynamics: an initial alignment phase that pushes toward the rank-1 signal and a secondary estimation phase where the dynamics approximate a one-dimensional SGD with controllable noise, benefiting from overparameterization to avoid trapping and improve generalization. This provides the first solid theoretical evidence that overparameterization can yield statistical advantages beyond exact parameterization in nonconvex tensor problems and points to potential extensions to other homogeneous models and online learning settings.
Abstract
We study the Order-$k$ ($k \geq 4$) spiked tensor model for the tensor principal component analysis (PCA) problem: given $N$ i.i.d. observations of a $k$-th order tensor generated from the model $\mathbf{T} = λ\cdot v_*^{\otimes k} + \mathbf{E}$, where $λ> 0$ is the signal-to-noise ratio (SNR), $v_*$ is a unit vector, and $\mathbf{E}$ is a random noise tensor, the goal is to recover the planted vector $v_*$. We propose a normalized stochastic gradient ascent (NSGA) method with overparameterization for solving the tensor PCA problem. Without any global (or spectral) initialization step, the proposed algorithm successfully recovers the signal $v_*$ when $Nλ^2 \geq \widetildeΩ(d^{\lceil k/2 \rceil})$, thereby breaking the previous conjecture that (stochastic) gradient methods require at least $Ω(d^{k-1})$ samples for recovery. For even $k$, the $\widetildeΩ(d^{k/2})$ threshold coincides with the optimal threshold under computational constraints, attained by sum-of-squares relaxations and related algorithms. Theoretical analysis demonstrates that the overparameterized stochastic gradient method not only establishes a significant initial optimization advantage during the early learning phase but also achieves strong generalization guarantees. This work provides the first evidence that overparameterization improves statistical performance relative to exact parameterization that is solved via standard continuous optimization.
