Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation
Ruchi Sandilya, Sumaira Perez, Charles Lynch, Lindsay Victoria, Benjamin Zebley, Derrick Matthew Buchanan, Mahendra T. Bhati, Nolan Williams, Timothy J. Spellman, Faith M. Gunning, Conor Liston, Logan Grosenick
TL;DR
Diffusion models offer high-fidelity generation but lack explicit dynamics-aware latent structure. ConDA introduces a dual-space framework that learns a compact contrastive embedding $\mathcal{C}$ of diffusion latents $\mathcal{Z}$ using auxiliary variables (e.g., time, stimulation), enabling nonlinear trajectory editing in $\mathcal{C}$ while rendering in $\mathcal{Z}$ for fidelity. The method combines a conditional latent diffusion model in $\mathcal{Z}$ with a contrastively learned, dynamics-aligned $\mathcal{C}$; editors such as spline and Taylor extrapolation enable interpolation, extrapolation, and class transfers with faithful reconstructions. Across five spatiotemporal domains (fluid dynamics, neural calcium imaging, therapeutic neurostimulation, facial expressions, and motor control), ConDA yields superior controllability and temporal coherence compared to linear traversals and conditioning baselines, demonstrating that diffusion latents encode dynamics-relevant structure that can be exploited via latent organization and manifold traversal for practical controllable generation.
Abstract
Diffusion models excel at generation, but their latent spaces are not explicitly organized for interpretable control. We introduce ConDA (Contrastive Diffusion Alignment), a framework that applies contrastive learning within diffusion embeddings to align latent geometry with system dynamics. Motivated by recent advances showing that contrastive objectives can recover more disentangled and structured representations, ConDA organizes diffusion latents such that traversal directions reflect underlying dynamical factors. Within this contrastively structured space, ConDA enables nonlinear trajectory traversal that supports faithful interpolation, extrapolation, and controllable generation. Across benchmarks in fluid dynamics, neural calcium imaging, therapeutic neurostimulation, and facial expression, ConDA produces interpretable latent representations with improved controllability compared to linear traversals and conditioning-based baselines. These results suggest that diffusion latents encode dynamics-relevant structure, but exploiting this structure requires latent organization and traversal along the latent manifold.
