CNN-Transformer Rectified Collaborative Learning for Medical Image Segmentation

Lanhu Wu; Miao Zhang; Yongri Piao; Zhenyan Yao; Weibing Sun; Feng Tian; Huchuan Lu

CNN-Transformer Rectified Collaborative Learning for Medical Image Segmentation

Lanhu Wu, Miao Zhang, Yongri Piao, Zhenyan Yao, Weibing Sun, Feng Tian, Huchuan Lu

TL;DR

This work tackles the mislocalization of CNN-based MIS and the coarse boundaries of Transformer-based MIS by enabling bi-directional knowledge transfer between CNN and Transformer models. It introduces Rectified Logit-wise Collaborative Learning (RLCL) to adaptively rectify wrong regions in soft labels using a ground-truth guided Adaptive Rectification Module (ARM) and Class-aware Feature-wise Collaborative Learning (CFCL) to align class-aware feature representations via a Category Perception Module (CPM). The framework optimizes segmentation jointly with logit- and feature-space distillation, using an overall loss that combines ${\mathcal{L}}_{seg}$, ${\mathcal{L}}_{rlcl}$, and ${\mathcal{L}}_{cfcl}$ terms, with dynamic weights $\lambda$ computed from alignment, similarity, and certainty factors $({\lambda}^a,{\lambda}^s,{\lambda}^c)$. Experiments on Synapse, ACDC, and Kvasir-SEG show CTRCL achieves new state-of-the-art results across multiple backbones, with notable improvements in DSC/JAC and reductions in HSD/MAE, and ablation studies confirm the benefits of both RLCL and CFCL and their robustness across architectures. The method’s generalization to diverse student pairs, offline KD, and existing MIS models suggests strong practical impact for improving medical image segmentation without prohibitive parameter growth.

Abstract

Automatic and precise medical image segmentation (MIS) is of vital importance for clinical diagnosis and analysis. Current MIS methods mainly rely on the convolutional neural network (CNN) or self-attention mechanism (Transformer) for feature modeling. However, CNN-based methods suffer from the inaccurate localization owing to the limited global dependency while Transformer-based methods always present the coarse boundary for the lack of local emphasis. Although some CNN-Transformer hybrid methods are designed to synthesize the complementary local and global information for better performance, the combination of CNN and Transformer introduces numerous parameters and increases the computation cost. To this end, this paper proposes a CNN-Transformer rectified collaborative learning (CTRCL) framework to learn stronger CNN-based and Transformer-based models for MIS tasks via the bi-directional knowledge transfer between them. Specifically, we propose a rectified logit-wise collaborative learning (RLCL) strategy which introduces the ground truth to adaptively select and rectify the wrong regions in student soft labels for accurate knowledge transfer in the logit space. We also propose a class-aware feature-wise collaborative learning (CFCL) strategy to achieve effective knowledge transfer between CNN-based and Transformer-based models in the feature space by granting their intermediate features the similar capability of category perception. Extensive experiments on three popular MIS benchmarks demonstrate that our CTRCL outperforms most state-of-the-art collaborative learning methods under different evaluation metrics.

CNN-Transformer Rectified Collaborative Learning for Medical Image Segmentation

TL;DR

, and

terms, with dynamic weights

computed from alignment, similarity, and certainty factors

. Experiments on Synapse, ACDC, and Kvasir-SEG show CTRCL achieves new state-of-the-art results across multiple backbones, with notable improvements in DSC/JAC and reductions in HSD/MAE, and ablation studies confirm the benefits of both RLCL and CFCL and their robustness across architectures. The method’s generalization to diverse student pairs, offline KD, and existing MIS models suggests strong practical impact for improving medical image segmentation without prohibitive parameter growth.

Abstract

Paper Structure (34 sections, 17 equations, 11 figures, 7 tables, 1 algorithm)

This paper contains 34 sections, 17 equations, 11 figures, 7 tables, 1 algorithm.

Introduction
Related Work
Medical Image Segmentation
Collaborative Learning
Proposed Method
Overview
Rectified Logit-wise Collaborative Learning
Alignment Factor (${\lambda}^{a}$)
Similarity-based Decay Factor (${\lambda}^{s}$)
Certainty-based Decay Factor (${\lambda}^{c}$)
Class-aware Feature-wise Collaborative Learning
Optimization
Experiments
Datasets
Synapse Multi-organ
...and 19 more sections

Figures (11)

Figure 1: Visualizations of class activation maps generated by Grad-CAM selvaraju2017grad and segmentation results of ResNet-50 he2016deep (Left) and MiT-B2 xie2021segformer (Right). Current CNN-based models suffer from the inaccurate localization (e.g., missing stomach in (a), (c)) while Transformer-based models present the coarse boundary (e.g., incomplete liver in (e), (g)). Our CTRCL improves the performance of the CNN-based model with more accurate location ((b), (d)) and the Transformer-based model with more elaborate boundary ((f), (h)) via collaborative learning between CNN-based and Transformer-based models.
Figure 2: The whole pipeline of our CTRCL framework, containing three parts: CNN-based student, Transformer-based student, and collaborative learning strategies (CFCL and RLCL). (a) Class-aware feature-wise collaborative learning (CFCL) focuses on the effective feature-wise knowledge transfer by encouraging student features to possess similar class-aware representations. (b) Rectified logit-wise collaborative learning (RLCL) aims at the accurate logit-wise knowledge transfer with student soft labels rectified by the ground truth. Please refer to Section \ref{['Methodology']} for details.
Figure 3: Illustration of the adaptive rectification module (ARM).
Figure 4: Illustration of the category perception module (CPM).
Figure 5: Visual comparisons on Synapse Multi-organ dataset. (a) MobileNetV2. (b) MiT-B0. (c) ResNet-50. (d) MiT-B2.
...and 6 more figures

CNN-Transformer Rectified Collaborative Learning for Medical Image Segmentation

TL;DR

Abstract

CNN-Transformer Rectified Collaborative Learning for Medical Image Segmentation

Authors

TL;DR

Abstract

Table of Contents

Figures (11)