Near-Optimality of Contrastive Divergence Algorithms
Pierre Glaser, Kevin Han Huang, Arthur Gretton
TL;DR
This work provides a non-asymptotic analysis of Contrastive Divergence (CD) for unnormalized exponential families, showing that online CD can attain the parametric rate $O(n^{-1/2})$ under mild regularity and sufficient MCMC steps, with averaging yielding near-Cramér-Rao optimal asymptotic variance. It further sharpens offline-CD bounds under subexponential tails, achieving near-parametric rates $O(( ext{log} n)^{1/2} n^{-1/2})$ by carefully controlling data-iterate correlations and tail probabilities, and demonstrates that Polyak-Ruppert averaging can yield estimators whose asymptotic variance is within a constant factor of the Fisher information bound. Across online and offline settings, the results establish the near-optimal statistical efficiency of CD-based training for unnormalized exponential families in the long-run regime, while detailing how batching, data reuse, and mixing assumptions influence convergence. The theoretical findings extend prior asymptotic results and provide practical guidance for achieving near-optimal performance with CD in finite-sample regimes.
Abstract
We perform a non-asymptotic analysis of the contrastive divergence (CD) algorithm, a training method for unnormalized models. While prior work has established that (for exponential family distributions) the CD iterates asymptotically converge at an $O(n^{-1 / 3})$ rate to the true parameter of the data distribution, we show, under some regularity assumptions, that CD can achieve the parametric rate $O(n^{-1 / 2})$. Our analysis provides results for various data batching schemes, including the fully online and minibatch ones. We additionally show that CD can be near-optimal, in the sense that its asymptotic variance is close to the Cramér-Rao lower bound.
