A Text-Image Fusion Method with Data Augmentation Capabilities for Referring Medical Image Segmentation
Shurong Chai, Rahul Kumar JAIN, Rui Xu, Shaocong Mo, Ruibo Hou, Shiyu Teng, Jiaqing Liu, Lanfen Lin, Yen-Wei Chen
TL;DR
This paper tackles the problem of aligning text descriptions with spatial regions in medical image segmentation when applying data augmentation. It introduces an early fusion framework that combines text features and visual information before augmentation, augmented by a ROI-guided lightweight text generator that produces a pseudo image to bridge semantic gaps. The model optimizes a joint loss $\mathcal{L}_{total} = \mathcal{L}_{Dice} + \lambda \mathcal{L}_{1}$ with $\lambda = 0.1$, enabling effective learning under diverse augmentations and achieving state-of-the-art performance across three datasets and four backbones. The approach offers practical value for clinical settings by preserving spatial consistency while leveraging diverse augmentations, and code is released on GitHub for reproducibility.
Abstract
Deep learning relies heavily on data augmentation to mitigate limited data, especially in medical imaging. Recent multimodal learning integrates text and images for segmentation, known as referring or text-guided image segmentation. However, common augmentations like rotation and flipping disrupt spatial alignment between image and text, weakening performance. To address this, we propose an early fusion framework that combines text and visual features before augmentation, preserving spatial consistency. We also design a lightweight generator that projects text embeddings into visual space, bridging semantic gaps. Visualization of generated pseudo-images shows accurate region localization. Our method is evaluated on three medical imaging tasks and four segmentation frameworks, achieving state-of-the-art results. Code is publicly available on GitHub: https://github.com/11yxk/MedSeg_EarlyFusion.
