Table of Contents
Fetching ...

Shrinkage to Infinity: Reducing Test Error by Inflating the Minimum Norm Interpolator in Linear Models

Jake Freeman

TL;DR

This work analyzes over-parameterized linear regression under strongly anisotropic covariances in the diverging $d/n$ regime. It shows that inflating the minimum-norm interpolator $\theta_{MN}$ by a factor $c>1$, i.e., using $c\theta_{MN}$, can reduce the generalization error $G(\cdot)$ when $\beta$ is aligned with top eigen-directions of $\Sigma$, a phenomenon termed the Inflation Property. The paper establishes both additive and multiplicative improvements, with the latter requiring stronger eigenstructure assumptions and yielding a near-optimal inflation constant $c_{opt}$ that can be approximated by $c_{opt}\approx \frac{\frac{n}{d}\beta^T\Sigma^2\beta}{\frac{n^2}{d^2}\beta^T\Sigma^3\beta}$ under suitable scaling. A data-splitting approach constructs a consistent estimator $\theta_{ds}$ whose inflation $c^*$ achieves $G(c^*\theta_{ds})=o(1)$ in probability, linking Johnson–Lindenstrauss-type projections to anti-regularization phenomena. These results rely on novel Gaussian-projection bounds for anisotropic covariances, and they illuminate how implicit regularization (and even anti-regularization) can improve prediction in high-dimensional, highly structured settings.

Abstract

Hastie et al. (2022) found that ridge regularization is essential in high dimensional linear regression $y=β^Tx + ε$ with isotropic co-variates $x\in \mathbb{R}^d$ and $n$ samples at fixed $d/n$. However, Hastie et al. (2022) also notes that when the co-variates are anisotropic and $β$ is aligned with the top eigenvalues of population covariance, the "situation is qualitatively different." In the present article, we make precise this observation for linear regression with highly anisotropic covariances and diverging $d/n$. We find that simply scaling up (or inflating) the minimum $\ell_2$ norm interpolator by a constant greater than one can improve the generalization error. This is in sharp contrast to traditional regularization/shrinkage prescriptions. Moreover, we use a data-splitting technique to produce consistent estimators that achieve generalization error comparable to that of the optimally inflated minimum-norm interpolator. Our proof relies on apparently novel matching upper and lower bounds for expectations of Gaussian random projections for a general class of anisotropic covariance matrices when $d/n\to \infty$.

Shrinkage to Infinity: Reducing Test Error by Inflating the Minimum Norm Interpolator in Linear Models

TL;DR

This work analyzes over-parameterized linear regression under strongly anisotropic covariances in the diverging regime. It shows that inflating the minimum-norm interpolator by a factor , i.e., using , can reduce the generalization error when is aligned with top eigen-directions of , a phenomenon termed the Inflation Property. The paper establishes both additive and multiplicative improvements, with the latter requiring stronger eigenstructure assumptions and yielding a near-optimal inflation constant that can be approximated by under suitable scaling. A data-splitting approach constructs a consistent estimator whose inflation achieves in probability, linking Johnson–Lindenstrauss-type projections to anti-regularization phenomena. These results rely on novel Gaussian-projection bounds for anisotropic covariances, and they illuminate how implicit regularization (and even anti-regularization) can improve prediction in high-dimensional, highly structured settings.

Abstract

Hastie et al. (2022) found that ridge regularization is essential in high dimensional linear regression with isotropic co-variates and samples at fixed . However, Hastie et al. (2022) also notes that when the co-variates are anisotropic and is aligned with the top eigenvalues of population covariance, the "situation is qualitatively different." In the present article, we make precise this observation for linear regression with highly anisotropic covariances and diverging . We find that simply scaling up (or inflating) the minimum norm interpolator by a constant greater than one can improve the generalization error. This is in sharp contrast to traditional regularization/shrinkage prescriptions. Moreover, we use a data-splitting technique to produce consistent estimators that achieve generalization error comparable to that of the optimally inflated minimum-norm interpolator. Our proof relies on apparently novel matching upper and lower bounds for expectations of Gaussian random projections for a general class of anisotropic covariance matrices when .
Paper Structure (14 sections, 64 theorems, 322 equations, 1 figure)

This paper contains 14 sections, 64 theorems, 322 equations, 1 figure.

Key Result

Theorem 2.8

Under Assumption assump:additive, there exists a universal constant $C>0$ such that for any $\frac{n}{d}\lambda_1\leq C$ and $n\gg 1$, Further, when $\lambda_1=o(\frac{d}{n})$ and $\sigma_{max}^2=\sigma^2$, Assumption assump:additive.3 is also a necessary condition for Equ. equ:additive_gen_err_1 to hold and the universal constant condition can be dropped.

Figures (1)

  • Figure 1: This figure illustrates how the Inflation Property relates to $\theta_{MN}$, positive ridge regression, and negative ridge regression.

Theorems & Definitions (123)

  • Definition 2.1: Linear Regression Setup
  • Definition 2.2: Minimum-Norm Interpolator
  • Definition 2.3: Generalization Error
  • Definition 2.4: Inflation Property
  • Definition 2.5
  • Definition 2.6
  • Theorem 2.8: Additive Improvement
  • Example 2.10
  • Theorem 2.12
  • Proposition 2.13
  • ...and 113 more