Table of Contents
Fetching ...

Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning

Dechen Zhang, Zhenmei Shi, Yi Zhang, Yingyu Liang, Difan Zou

TL;DR

This work develops kernel ridge regression theory for structured non-i.i.d. data generated by a signal-noise causal mechanism, introducing a blockwise decomposition that yields a Bernstein-type concentration bound for $k$-gap independent observations. It derives excess risk bounds $R(\lambda)$ that decompose into a bias term $\mathrm{Bias}^2(\lambda)$ and a variance term $\mathrm{Var}(\lambda)$, with rates governed by the kernel spectrum decay, smoothness of the target, and sampling parameters including the signal-to-noise relevance $\tilde{r}$ and the noise multiplicity $k$. Under conditional orthogonality, the bound sharpens by replacing $r_T$ with $r_0$ and $r_e$, refining how data dependence affects generalization. The authors then apply the framework to denoising score learning and DDPMs, providing guidance on optimal noise multiplicity $k$ as a function of the time step and noise level, and corroborate the theory with numerical experiments showing the predicted trade-offs. Overall, the paper advances KRR generalization theory in dependent-data regimes and offers principled sampling strategies for denoising and diffusion-model training.

Abstract

Kernel ridge regression (KRR) is a foundational tool in machine learning, with recent work emphasizing its connections to neural networks. However, existing theory primarily addresses the i.i.d. setting, while real-world data often exhibits structured dependencies - particularly in applications like denoising score learning where multiple noisy observations derive from shared underlying signals. We present the first systematic study of KRR generalization for non-i.i.d. data with signal-noise causal structure, where observations represent different noisy views of common signals. By developing a novel blockwise decomposition method that enables precise concentration analysis for dependent data, we derive excess risk bounds for KRR that explicitly depend on: (1) the kernel spectrum, (2) causal structure parameters, and (3) sampling mechanisms (including relative sample sizes for signals and noises). We further apply our results to denoising score learning, establishing generalization guarantees and providing principled guidance for sampling noisy data points. This work advances KRR theory while providing practical tools for analyzing dependent data in modern machine learning applications.

Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning

TL;DR

This work develops kernel ridge regression theory for structured non-i.i.d. data generated by a signal-noise causal mechanism, introducing a blockwise decomposition that yields a Bernstein-type concentration bound for -gap independent observations. It derives excess risk bounds that decompose into a bias term and a variance term , with rates governed by the kernel spectrum decay, smoothness of the target, and sampling parameters including the signal-to-noise relevance and the noise multiplicity . Under conditional orthogonality, the bound sharpens by replacing with and , refining how data dependence affects generalization. The authors then apply the framework to denoising score learning and DDPMs, providing guidance on optimal noise multiplicity as a function of the time step and noise level, and corroborate the theory with numerical experiments showing the predicted trade-offs. Overall, the paper advances KRR generalization theory in dependent-data regimes and offers principled sampling strategies for denoising and diffusion-model training.

Abstract

Kernel ridge regression (KRR) is a foundational tool in machine learning, with recent work emphasizing its connections to neural networks. However, existing theory primarily addresses the i.i.d. setting, while real-world data often exhibits structured dependencies - particularly in applications like denoising score learning where multiple noisy observations derive from shared underlying signals. We present the first systematic study of KRR generalization for non-i.i.d. data with signal-noise causal structure, where observations represent different noisy views of common signals. By developing a novel blockwise decomposition method that enables precise concentration analysis for dependent data, we derive excess risk bounds for KRR that explicitly depend on: (1) the kernel spectrum, (2) causal structure parameters, and (3) sampling mechanisms (including relative sample sizes for signals and noises). We further apply our results to denoising score learning, establishing generalization guarantees and providing principled guidance for sampling noisy data points. This work advances KRR theory while providing practical tools for analyzing dependent data in modern machine learning applications.
Paper Structure (38 sections, 43 theorems, 344 equations, 3 figures)

This paper contains 38 sections, 43 theorems, 344 equations, 3 figures.

Key Result

Theorem 1

(Informal statement of Theorem main theorem) Under general assumptions, if the regularization parameter $\lambda=\Omega\left(n^{-\beta}\right)$, then the asymptotic rate (with respect to sample size $n$) of the generalization error (excess risk) $R(\lambda)$ is roughly where $\beta$ denotes the decay rate of the kernel eigenvalues, $\Tilde{s}$ represents the smoothness of target function, $\Tilde

Figures (3)

  • Figure 1: Causal structure of our data model.
  • Figure 2: Score estimation error (mean $\pm$ s.d.) versus the number of noise per data, i.e., $k$, for four noise levels, where lower error implies better score learning.
  • Figure 3: Score estimation error versus the number of noise per data, i.e., $k$, for two noise levels.

Theorems & Definitions (77)

  • Theorem 1
  • Example 3.1
  • Example 3.2
  • Definition 3.1
  • Definition 3.2
  • Definition 3.3
  • Theorem 4.1
  • Remark 4.2
  • Theorem 4.3
  • Theorem 4.4
  • ...and 67 more