FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning
Ningxin He, Yang Liu, Wei Sun, Xiaozhou Ye, Ye Ouyang, Tiegang Gao, Zehui Zhang
TL;DR
FedMMKT introduces a privacy-preserving multimodal federated learning framework that co-enhances a server text-to-image (T2I) model and client task models using synthetic cross-modal data generated on the server. It employs LabVote label refinement and MultiRepFusion multimodal representation fusion to robustly align knowledge from heterogeneous clients, enabling bidirectional knowledge transfer without sharing raw data or model parameters. Theoretical convergence guarantees under non-IID data and multi-modal alignment errors are provided, and experiments on Flowers102 and Food101 show substantial gains for both clients and the server along with favorable communication costs. Overall, FedMMKT demonstrates a practical, scalable approach to cross-silo knowledge transfer in privacy-sensitive, multimodal FL settings.
Abstract
Text-to-Image (T2I) models have demonstrated their versatility in a wide range of applications. However, adaptation of T2I models to specialized tasks is often limited by the availability of task-specific data due to privacy concerns. On the other hand, harnessing the power of rich multimodal data from modern mobile systems and IoT infrastructures presents a great opportunity. This paper introduces Federated Multi-modal Knowledge Transfer (FedMMKT), a novel framework that enables co-enhancement of a server T2I model and client task-specific models using decentralized multimodal data without compromising data privacy.
