Multimodal Datasets with Controllable Mutual Information
Raheem Karim Hashmani, Garrett W. Merz, Helen Qu, Mariel Pettee, Kyle Cranmer
TL;DR
This work tackles the challenge of understanding and benchmarking how mutual information between multiple modalities influences multimodal self-supervised learning. It proposes a three-stage framework $\mathbf{u} \to \mathbf{z} \to \mathbf{x}$ that uses a linear causal construction to produce latent variables with known MI, then applies invertible flow-based mappings to generate realistic multimodal data while preserving that MI. The key contributions are: (i) analytic or semi-analytic MI expressions for latent variables, (ii) a flow-based MI-preserving data-generation pipeline, (iii) templates to control how information is distributed across modalities for targeted ablations, and (iv) concrete dataset examples including a black-hole imaging scenario and scalable massively multimodal data, with reproducible code. The resulting datasets provide a principled, controllable testbed for MI estimators and for studying the role of information overlap in multimodal SSL as the number of modalities grows, with practical impact on benchmarking and model design.
Abstract
We introduce a framework for generating highly multimodal datasets with explicitly calculable mutual information between modalities. This enables the construction of benchmark datasets that provide a novel testbed for systematic studies of mutual information estimators and multimodal self-supervised learning techniques. Our framework constructs realistic datasets with known mutual information using a flow-based generative model and a structured causal framework for generating correlated latent variables.
