Table of Contents
Fetching ...

Particle Dynamics for Latent-Variable Energy-Based Models

Shiqin Tang, Shuxin Zhuang, Rong Feng, Runsheng Yu, Hongzong Li, Youzhi Zhang

TL;DR

This work tackles learning latent-variable energy-based models (LV-EBMs) by reframing maximum-likelihood estimation as a saddle problem over distributions on the latent and joint data manifolds. It introduces a particle-based, energy-driven algorithm built on coupled Wasserstein gradient flows, realized via overdamped Langevin dynamics for both conditional latent variables and joint samples, with stochastic ascent for the energy parameters $\theta$ and no reliance on discriminators or decoders. The authors establish well-posedness and convergence of the inner flows under standard regularity (e.g., log-Sobolev inequalities) and show a variational-training variant yields a tighter ELBO than traditional VI bounds. Empirically, LV-EBMs demonstrate competitive or superior performance to VAE, non-amortized VI, and Hard EM on controlled synthetic geometries and real-world UCI data, with improved likelihood-like objectives and faithful reconstruction. The framework offers a principled, scalable path toward energy-based latent-variable modeling that preserves multimodality and conditional structure while avoiding amortized inference biases.

Abstract

Latent-variable energy-based models (LVEBMs) assign a single normalized energy to joint pairs of observed data and latent variables, offering expressive generative modeling while capturing hidden structure. We recast maximum-likelihood training as a saddle problem over distributions on the latent and joint manifolds and view the inner updates as coupled Wasserstein gradient flows. The resulting algorithm alternates overdamped Langevin updates for a joint negative pool and for conditional latent particles with stochastic parameter ascent, requiring no discriminator or auxiliary networks. We prove existence and convergence under standard smoothness and dissipativity assumptions, with decay rates in KL divergence and Wasserstein-2 distance. The saddle-point view further yields an ELBO strictly tighter than bounds obtained with restricted amortized posteriors. Our method is evaluated on numerical approximations of physical systems and performs competitively against comparable approaches.

Particle Dynamics for Latent-Variable Energy-Based Models

TL;DR

This work tackles learning latent-variable energy-based models (LV-EBMs) by reframing maximum-likelihood estimation as a saddle problem over distributions on the latent and joint data manifolds. It introduces a particle-based, energy-driven algorithm built on coupled Wasserstein gradient flows, realized via overdamped Langevin dynamics for both conditional latent variables and joint samples, with stochastic ascent for the energy parameters and no reliance on discriminators or decoders. The authors establish well-posedness and convergence of the inner flows under standard regularity (e.g., log-Sobolev inequalities) and show a variational-training variant yields a tighter ELBO than traditional VI bounds. Empirically, LV-EBMs demonstrate competitive or superior performance to VAE, non-amortized VI, and Hard EM on controlled synthetic geometries and real-world UCI data, with improved likelihood-like objectives and faithful reconstruction. The framework offers a principled, scalable path toward energy-based latent-variable modeling that preserves multimodality and conditional structure while avoiding amortized inference biases.

Abstract

Latent-variable energy-based models (LVEBMs) assign a single normalized energy to joint pairs of observed data and latent variables, offering expressive generative modeling while capturing hidden structure. We recast maximum-likelihood training as a saddle problem over distributions on the latent and joint manifolds and view the inner updates as coupled Wasserstein gradient flows. The resulting algorithm alternates overdamped Langevin updates for a joint negative pool and for conditional latent particles with stochastic parameter ascent, requiring no discriminator or auxiliary networks. We prove existence and convergence under standard smoothness and dissipativity assumptions, with decay rates in KL divergence and Wasserstein-2 distance. The saddle-point view further yields an ELBO strictly tighter than bounds obtained with restricted amortized posteriors. Our method is evaluated on numerical approximations of physical systems and performs competitively against comparable approaches.
Paper Structure (34 sections, 4 theorems, 35 equations, 3 figures, 3 tables, 1 algorithm)

This paper contains 34 sections, 4 theorems, 35 equations, 3 figures, 3 tables, 1 algorithm.

Key Result

Proposition 1

Under Assumption 1, for any $\tilde{q}\in\tilde{\mathcal{Q}}$ and any $x\in\mathbb{R}^d$, $q \in\mathcal{Q}$, Hence the inner optima in eq:main_eq are and in each case the minimizer is unique up to $\nu\otimes\mu$ (resp. $\mu$) null sets.

Figures (3)

  • Figure 1: Training loss vs. step (up to 4500) for three scenarios. Solid lines depict the mean over three runs; shaded regions show ±1 standard deviation.
  • Figure 2: Left: Ten snapshots of LCR-2D reconstructions across training steps; Colors denote ground-truth radial classes. Right: Radius–step heatmap computed from the same reconstructions. The color intensity encodes the per-step normalized radial density.
  • Figure 3: Reconstruction results: LCS-3D (top), LCR-2D (bottom-left), and HMR-2D (bottom-right).

Theorems & Definitions (9)

  • Remark 1
  • Proposition 1: Inner optima for \ref{['eq:main_eq']}: joint and conditional Gibbs principles
  • proof
  • Theorem 1: Convergence of inner Fokker–Planck flows
  • proof : Proof sketch
  • Theorem 2
  • proof
  • Theorem 3: Weak convergence without LSI
  • proof