Table of Contents
Fetching ...

Overparametrization bends the landscape: BBP transitions at initialization in simple Neural Networks

Brandon Livio Annesi, Dario Bocchi, Chiara Cammarota

TL;DR

This work analyzes BBP-type transitions in the Hessian spectrum of a teacher–student two-layer soft-committee machine with quadratic activations, focusing on the spectrum at initialization in the high-dimensional limit. Using a field-theoretic approach, the authors derive self-consistent equations for the bulk and outlier eigenvalues and show that overparameterization shifts the BBP transition to lower data requirements and can change its nature from continuous to discontinuous. In the large overparameterization limit ($p o ty$), the BBP threshold becomes $oldsymbol{ extalpha}_{ ext{BBP}}^{p= ty} = rac{p^*(a+1)}{2}$, with the minimum at $a=0$ recovering the information-theoretic weak-recovery bound $oldsymbol{ extalpha}= rac{p^*}{2}$. Numerical simulations at finite $N$ reveal strong finite-size effects for discontinuous transitions and show how, despite these corrections, overparameterization generally facilitates earlier information emergence in the Hessian, supporting spectral initialization benefits. Overall, the work connects spectral properties of random Hessians to learning performance, highlighting how overparameterization reshapes optimization landscapes and informs initialization-based strategies for efficient learning.

Abstract

High-dimensional non-convex loss landscapes play a central role in the theory of Machine Learning. Gaining insight into how these landscapes interact with gradient-based optimization methods, even in relatively simple models, can shed light on this enigmatic feature of neural networks. In this work, we will focus on a prototypical simple learning problem, which generalizes the Phase Retrieval inference problem by allowing the exploration of overparametrized settings. Using techniques from field theory, we analyze the spectrum of the Hessian at initialization and identify a Baik-Ben Arous-Péché (BBP) transition in the amount of data that separates regimes where the initialization is informative or uninformative about a planted signal of a teacher-student setup. Crucially, we demonstrate how overparameterization can bend the loss landscape, shifting the transition point, even reaching the information-theoretic weak-recovery threshold in the large overparameterization limit, while also altering its qualitative nature. We distinguish between continuous and discontinuous BBP transitions and support our analytical predictions with simulations, examining how they compare to the finite-N behavior. In the case of discontinuous BBP transitions strong finite-N corrections allow the retrieval of information at a signal-to-noise ratio (SNR) smaller than the predicted BBP transition. In these cases we provide estimates for a new lower SNR threshold that marks the point at which initialization becomes entirely uninformative.

Overparametrization bends the landscape: BBP transitions at initialization in simple Neural Networks

TL;DR

This work analyzes BBP-type transitions in the Hessian spectrum of a teacher–student two-layer soft-committee machine with quadratic activations, focusing on the spectrum at initialization in the high-dimensional limit. Using a field-theoretic approach, the authors derive self-consistent equations for the bulk and outlier eigenvalues and show that overparameterization shifts the BBP transition to lower data requirements and can change its nature from continuous to discontinuous. In the large overparameterization limit (), the BBP threshold becomes , with the minimum at recovering the information-theoretic weak-recovery bound . Numerical simulations at finite reveal strong finite-size effects for discontinuous transitions and show how, despite these corrections, overparameterization generally facilitates earlier information emergence in the Hessian, supporting spectral initialization benefits. Overall, the work connects spectral properties of random Hessians to learning performance, highlighting how overparameterization reshapes optimization landscapes and informs initialization-based strategies for efficient learning.

Abstract

High-dimensional non-convex loss landscapes play a central role in the theory of Machine Learning. Gaining insight into how these landscapes interact with gradient-based optimization methods, even in relatively simple models, can shed light on this enigmatic feature of neural networks. In this work, we will focus on a prototypical simple learning problem, which generalizes the Phase Retrieval inference problem by allowing the exploration of overparametrized settings. Using techniques from field theory, we analyze the spectrum of the Hessian at initialization and identify a Baik-Ben Arous-Péché (BBP) transition in the amount of data that separates regimes where the initialization is informative or uninformative about a planted signal of a teacher-student setup. Crucially, we demonstrate how overparameterization can bend the loss landscape, shifting the transition point, even reaching the information-theoretic weak-recovery threshold in the large overparameterization limit, while also altering its qualitative nature. We distinguish between continuous and discontinuous BBP transitions and support our analytical predictions with simulations, examining how they compare to the finite-N behavior. In the case of discontinuous BBP transitions strong finite-N corrections allow the retrieval of information at a signal-to-noise ratio (SNR) smaller than the predicted BBP transition. In these cases we provide estimates for a new lower SNR threshold that marks the point at which initialization becomes entirely uninformative.
Paper Structure (21 sections, 58 equations, 7 figures)

This paper contains 21 sections, 58 equations, 7 figures.

Figures (7)

  • Figure 1: Overlap between the signal estimate and the true signal as a function of $\alpha$ for continuous BBP (Left) and discontinuous BBP (Right), with $p=2$ and $p^*=1$.
  • Figure 2: $\alpha_{BBP}$ as a function of $a$ for $p^* = 1$ and several values of $p$. The point at which the curves start increasing almost linearly is the point in which the transition becomes discontinuous. The dashed line shows $\alpha_{BBP}(a)$ in the large overparametrization limit where the transition is always discontinuous. The insets show $\alpha_{BBP}$ as a function of $p$ for three fixed values of $a$. Here, red points indicate the transition is continuous, while blue points that it is discontinuous. The red crosses are estimates of $\alpha_0$, a "finite-$N$" estimate of the transition described in section \ref{['sec:numericalBBP']}.
  • Figure 3: Comparison of BBP transitions for $p = p^* =1$. On the $y$ axis we plot $\phi$, defined as the fraction of times the eigenvector with the maximum overlap with the signal corresponds to the smallest eigenvalue. On the left a value of $a$ for which the transition is continuous, on the right a value for which it is discontinuous. The vertical blue lines show our prediction for the BBP threshold, while for the discontinuous case the red line shows our estimate of $\alpha_{0}$.
  • Figure 4: $\alpha_{BBP}$ (solid) and $\alpha_0$ (crosses) for two different values of $p$ as a function of $a$.
  • Figure 5: In Panel \ref{['left-figure']}, a continuous BBP is shown: the $\alpha_{\mathrm{BBP}}$ corresponds to the crossing of the $g^{*}(\alpha)$ and $g_{-}(\alpha)$ curves. In Panel \ref{['right-figure']}, a discontinuous BBP is shown: the $\alpha_{\mathrm{BBP}}$ corresponds to the crossing of the $g^{*}(\alpha)$ and $g_{\min}(\alpha)$ curves.
  • ...and 2 more figures