Table of Contents
Fetching ...

Near-Optimality of Contrastive Divergence Algorithms

Pierre Glaser, Kevin Han Huang, Arthur Gretton

TL;DR

This work provides a non-asymptotic analysis of Contrastive Divergence (CD) for unnormalized exponential families, showing that online CD can attain the parametric rate $O(n^{-1/2})$ under mild regularity and sufficient MCMC steps, with averaging yielding near-Cramér-Rao optimal asymptotic variance. It further sharpens offline-CD bounds under subexponential tails, achieving near-parametric rates $O(( ext{log} n)^{1/2} n^{-1/2})$ by carefully controlling data-iterate correlations and tail probabilities, and demonstrates that Polyak-Ruppert averaging can yield estimators whose asymptotic variance is within a constant factor of the Fisher information bound. Across online and offline settings, the results establish the near-optimal statistical efficiency of CD-based training for unnormalized exponential families in the long-run regime, while detailing how batching, data reuse, and mixing assumptions influence convergence. The theoretical findings extend prior asymptotic results and provide practical guidance for achieving near-optimal performance with CD in finite-sample regimes.

Abstract

We perform a non-asymptotic analysis of the contrastive divergence (CD) algorithm, a training method for unnormalized models. While prior work has established that (for exponential family distributions) the CD iterates asymptotically converge at an $O(n^{-1 / 3})$ rate to the true parameter of the data distribution, we show, under some regularity assumptions, that CD can achieve the parametric rate $O(n^{-1 / 2})$. Our analysis provides results for various data batching schemes, including the fully online and minibatch ones. We additionally show that CD can be near-optimal, in the sense that its asymptotic variance is close to the Cramér-Rao lower bound.

Near-Optimality of Contrastive Divergence Algorithms

TL;DR

This work provides a non-asymptotic analysis of Contrastive Divergence (CD) for unnormalized exponential families, showing that online CD can attain the parametric rate under mild regularity and sufficient MCMC steps, with averaging yielding near-Cramér-Rao optimal asymptotic variance. It further sharpens offline-CD bounds under subexponential tails, achieving near-parametric rates by carefully controlling data-iterate correlations and tail probabilities, and demonstrates that Polyak-Ruppert averaging can yield estimators whose asymptotic variance is within a constant factor of the Fisher information bound. Across online and offline settings, the results establish the near-optimal statistical efficiency of CD-based training for unnormalized exponential families in the long-run regime, while detailing how batching, data reuse, and mixing assumptions influence convergence. The theoretical findings extend prior asymptotic results and provide practical guidance for achieving near-optimal performance with CD in finite-sample regimes.

Abstract

We perform a non-asymptotic analysis of the contrastive divergence (CD) algorithm, a training method for unnormalized models. While prior work has established that (for exponential family distributions) the CD iterates asymptotically converge at an rate to the true parameter of the data distribution, we show, under some regularity assumptions, that CD can achieve the parametric rate . Our analysis provides results for various data batching schemes, including the fully online and minibatch ones. We additionally show that CD can be near-optimal, in the sense that its asymptotic variance is close to the Cramér-Rao lower bound.
Paper Structure (52 sections, 37 theorems, 241 equations, 1 figure, 2 algorithms)

This paper contains 52 sections, 37 theorems, 241 equations, 1 figure, 2 algorithms.

Key Result

Lemma 3.1

Let $(\psi_t)_{0 \leq t \leq n}$ be the iterates from Algorithm alg:online_cd. Denote $\delta_{t} = \mathbb{ E } \| \psi_{t} - {\psi}^{\star} \|^2$, $\sigma_{\star} = (\mathbb{ E } \| \phi(X_1) - \mathbb{ E } \left \lbrack \phi(X_1) \right \rbrack \|^2)^{1 / 2}$, and $\sigma_{t} = (\mathbb{ E } \| where $\left \| \log Z \right \|_{3, \infty}$ is a constant, $\tilde{\mu}_{m, t} \coloneqq \mu - \a

Figures (1)

  • Figure :

Theorems & Definitions (65)

  • Lemma 3.1
  • Theorem 3.2
  • Theorem 3.3: Contrastive Divergence with Polyak-Ruppert averaging
  • Theorem 4.1: Theorem 2.1 of jiang2018convergence
  • Theorem 4.2
  • Theorem 4.3: Convergence up to a tail control
  • Lemma 4.4
  • Theorem 4.5
  • Remark : Examples
  • Theorem B.1
  • ...and 55 more