SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment
Abdulmomen Ghalkha, Zhuojun Tian, Chaouki Ben Issaid, Mehdi Bennis
TL;DR
SheafAlign introduces a decentralized, shear-theoretic framework for multimodal alignment that abandons a single shared embedding space in favor of a network of pairwise comparison spaces defined on a communication graph. Each edge maintains a shared comparison space with restriction maps and a dual reconstruction mechanism, enabling robust alignment across modalities while preserving modality-specific information. The training objective combines a sheaf Laplacian consistency term, edge-wise InfoNCE contrastive losses, and cross-edge reconstruction losses, all optimized in a fully decentralized manner. Experiments on real and synthetic datasets demonstrate superior zero-shot and few-shot generalization, improved cross-modal retrieval, and robust performance under missing modalities, with up to 50% savings in communication cost compared to baselines like ImageBind. This approach offers a practical path to scalable, information-preserving multimodal fusion in distributed sensing environments.
Abstract
Conventional multimodal alignment methods assume mutual redundancy across all modalities, an assumption that fails in real-world distributed scenarios. We propose SheafAlign, a sheaf-theoretic framework for decentralized multimodal alignment that replaces single-space alignment with multiple comparison spaces. This approach models pairwise modality relations through sheaf structures and leverages decentralized contrastive learning-based objectives for training. SheafAlign overcomes the limitations of prior methods by not requiring mutual redundancy among all modalities, preserving both shared and unique information. Experiments on multimodal sensing datasets show superior zero-shot generalization, cross-modal alignment, and robustness to missing modalities, with 50\% lower communication cost than state-of-the-art baselines.
