Table of Contents
Fetching ...

Transfer Learning for Benign Overfitting in High-Dimensional Linear Regression

Yeichan Kim, Ilmun Kim, Seyoung Park

TL;DR

This research proposes a novel two-step Transfer MNI approach, characterize its non-asymptotic excess risk and identify conditions under which it outperforms the target-only MNI, and reveals free-lunch covariate shift regimes, where leveraging heterogeneous data yields the benefit of knowledge transfer at limited cost.

Abstract

Transfer learning is a key component of modern machine learning, enhancing the performance of target tasks by leveraging diverse data sources. Simultaneously, overparameterized models such as the minimum-$\ell_2$-norm interpolator (MNI) in high-dimensional linear regression have garnered significant attention for their remarkable generalization capabilities, a property known as benign overfitting. Despite their individual importance, the intersection of transfer learning and MNI remains largely unexplored. Our research bridges this gap by proposing a novel two-step Transfer MNI approach and analyzing its trade-offs. We characterize its non-asymptotic excess risk and identify conditions under which it outperforms the target-only MNI. Our analysis reveals free-lunch covariate shift regimes, where leveraging heterogeneous data yields the benefit of knowledge transfer at limited cost. To operationalize our findings, we develop a data-driven procedure to detect informative sources and introduce an ensemble method incorporating multiple informative Transfer MNIs. Finite-sample experiments demonstrate the robustness of our methods to model and data heterogeneity, confirming their advantage.

Transfer Learning for Benign Overfitting in High-Dimensional Linear Regression

TL;DR

This research proposes a novel two-step Transfer MNI approach, characterize its non-asymptotic excess risk and identify conditions under which it outperforms the target-only MNI, and reveals free-lunch covariate shift regimes, where leveraging heterogeneous data yields the benefit of knowledge transfer at limited cost.

Abstract

Transfer learning is a key component of modern machine learning, enhancing the performance of target tasks by leveraging diverse data sources. Simultaneously, overparameterized models such as the minimum--norm interpolator (MNI) in high-dimensional linear regression have garnered significant attention for their remarkable generalization capabilities, a property known as benign overfitting. Despite their individual importance, the intersection of transfer learning and MNI remains largely unexplored. Our research bridges this gap by proposing a novel two-step Transfer MNI approach and analyzing its trade-offs. We characterize its non-asymptotic excess risk and identify conditions under which it outperforms the target-only MNI. Our analysis reveals free-lunch covariate shift regimes, where leveraging heterogeneous data yields the benefit of knowledge transfer at limited cost. To operationalize our findings, we develop a data-driven procedure to detect informative sources and introduce an ensemble method incorporating multiple informative Transfer MNIs. Finite-sample experiments demonstrate the robustness of our methods to model and data heterogeneity, confirming their advantage.
Paper Structure (33 sections, 18 theorems, 187 equations, 4 figures, 1 table, 1 algorithm)

This paper contains 33 sections, 18 theorems, 187 equations, 4 figures, 1 table, 1 algorithm.

Key Result

Lemma 1

Under the mutual independence and mean-zero condition of $(\mathcal{X},\mathcal{E})$ in Assumption assumption:model, the excess risk of the TM estimate is the sum of bias $\mathcal{B}_{ \mathrm{TM}}^{(q)}$ and variance $\mathcal{V}_{ \mathrm{TM}}^{(q)}$ such that

Figures (4)

  • Figure 1: We $n_0=25$ and $n_1=n_2=n_3=75$ for overfitting with $S=500$. Figures (b) and (c) incorporate covariate shifts as detailed, with (c) additionally benefiting from the "free-lunch" effect.
  • Figure 2: We set $n_0=n_1=50$, and $n_2=\lfloor n_2^{\ast} \rfloor$, the optimal transfer size for $\hat{{\boldsymbol\beta}}_{ \mathrm{TM}}^{(2)}$, with $S=10$. Figure (c) adjusts $n_2^{\ast}$ in Corollary \ref{['cor:isotropic']} via the modified signal-to-noise ratio $\mathrm{SNR}_{\alpha} := \alpha \|{\boldsymbol\beta}^{(0)}\|^2/\sigma^2$, which allows $\hat{{\boldsymbol\beta}}_{ \mathrm{TM}}^{(2)}$ to leverage even more source samples than the original $\lfloor n_2^{\ast} \rfloor$ in (a) and (b).
  • Figure 3: We compare the performance of our proposed methods to single-source-pooled-MNIs under model shift, with the same setting as Figure \ref{['fig:benign']}. Figure (a) re-uses the exact seeds for Figure \ref{['fig:benign']}, so the excess risk curves for $\hat{{\boldsymbol\beta}}_{ \mathrm{M}}^{(0)}$, $\hat{{\boldsymbol\beta}}_{ \mathrm{TM}}^{(q)}$, and $\hat{{\boldsymbol\beta}}_{ \mathrm{WTM}}$ are identical across the two figures. In contrast, (b) uses distinct seeds, hence its curves differ from those in Figure \ref{['fig:benign']}.
  • Figure 4: As in Figure \ref{['fig:harmless']}, the target and source covariates are i.i.d. $\mathcal{N}(\mathbf{0}_p,\mathbf{I}_p)$ with $\mathrm{SNR}=10$. Once $n_1=n_1^{\ast}$ is determined by \ref{['eq:n1_star']}, the corresponding target sample size $n_0^{\ast}$ follows from \ref{['eq:n0_star']}.

Theorems & Definitions (33)

  • Definition 1
  • Definition 2
  • Lemma 1
  • Remark 1: Bias reduction versus variance inflation
  • Theorem 1
  • Corollary 1
  • Theorem 2
  • Corollary 2: Free-lunch covariate shift
  • Remark 2: Relaxed free-lunch condition
  • Lemma 2: Expectation of quadratic forms
  • ...and 23 more