Table of Contents
Fetching ...

SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment

Abdulmomen Ghalkha, Zhuojun Tian, Chaouki Ben Issaid, Mehdi Bennis

TL;DR

SheafAlign introduces a decentralized, shear-theoretic framework for multimodal alignment that abandons a single shared embedding space in favor of a network of pairwise comparison spaces defined on a communication graph. Each edge maintains a shared comparison space with restriction maps and a dual reconstruction mechanism, enabling robust alignment across modalities while preserving modality-specific information. The training objective combines a sheaf Laplacian consistency term, edge-wise InfoNCE contrastive losses, and cross-edge reconstruction losses, all optimized in a fully decentralized manner. Experiments on real and synthetic datasets demonstrate superior zero-shot and few-shot generalization, improved cross-modal retrieval, and robust performance under missing modalities, with up to 50% savings in communication cost compared to baselines like ImageBind. This approach offers a practical path to scalable, information-preserving multimodal fusion in distributed sensing environments.

Abstract

Conventional multimodal alignment methods assume mutual redundancy across all modalities, an assumption that fails in real-world distributed scenarios. We propose SheafAlign, a sheaf-theoretic framework for decentralized multimodal alignment that replaces single-space alignment with multiple comparison spaces. This approach models pairwise modality relations through sheaf structures and leverages decentralized contrastive learning-based objectives for training. SheafAlign overcomes the limitations of prior methods by not requiring mutual redundancy among all modalities, preserving both shared and unique information. Experiments on multimodal sensing datasets show superior zero-shot generalization, cross-modal alignment, and robustness to missing modalities, with 50\% lower communication cost than state-of-the-art baselines.

SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment

TL;DR

SheafAlign introduces a decentralized, shear-theoretic framework for multimodal alignment that abandons a single shared embedding space in favor of a network of pairwise comparison spaces defined on a communication graph. Each edge maintains a shared comparison space with restriction maps and a dual reconstruction mechanism, enabling robust alignment across modalities while preserving modality-specific information. The training objective combines a sheaf Laplacian consistency term, edge-wise InfoNCE contrastive losses, and cross-edge reconstruction losses, all optimized in a fully decentralized manner. Experiments on real and synthetic datasets demonstrate superior zero-shot and few-shot generalization, improved cross-modal retrieval, and robust performance under missing modalities, with up to 50% savings in communication cost compared to baselines like ImageBind. This approach offers a practical path to scalable, information-preserving multimodal fusion in distributed sensing environments.

Abstract

Conventional multimodal alignment methods assume mutual redundancy across all modalities, an assumption that fails in real-world distributed scenarios. We propose SheafAlign, a sheaf-theoretic framework for decentralized multimodal alignment that replaces single-space alignment with multiple comparison spaces. This approach models pairwise modality relations through sheaf structures and leverages decentralized contrastive learning-based objectives for training. SheafAlign overcomes the limitations of prior methods by not requiring mutual redundancy among all modalities, preserving both shared and unique information. Experiments on multimodal sensing datasets show superior zero-shot generalization, cross-modal alignment, and robustness to missing modalities, with 50\% lower communication cost than state-of-the-art baselines.
Paper Structure (9 sections, 6 equations, 3 figures, 1 table, 1 algorithm)

This paper contains 9 sections, 6 equations, 3 figures, 1 table, 1 algorithm.

Figures (3)

  • Figure 1: Bottom: Illustration of the absence of mutual redundancies between modalities. Top: Schematic diagram of SheafAlign architecture denoting restriction maps and comparison space alignment.
  • Figure 2: Zero-shot and few-shot test accuracy for SheafAlign compared to ImageBind and supervised baselines across (a) DeepSense blockage prediction, and (b) Multi-View MNIST datasets, with zero-shot accuracies denoted as $\star$ and $\blacktriangle$ for SheafAlign and ImageBind, respectively.
  • Figure 3: Cross-modal retrieval performance of SheafAlign compared to ImageBind across (a) Multi-view MNIST, and (b) semantic inpainting datasets. Bars show the mean Recall@K ($K = 1, 5, 10$), averaged across all modality pairs.

Theorems & Definitions (1)

  • Definition 1: Information-Theoretic Decomposition