SCHNet: SAM Marries CLIP for Human Parsing

Kunliang Liu; Jianming Wang; Rize Jin; Wonjun Hwang; Tae-Sun Chung

SCHNet: SAM Marries CLIP for Human Parsing

Kunliang Liu, Jianming Wang, Rize Jin, Wonjun Hwang, Tae-Sun Chung

TL;DR

This work tackles semantic-aware human parsing by integrating CLIP's semantic understanding with SAM's fine-grained segmentation. It introduces Semantic-Refinement Module (SRM) to inject multi-level CLIP semantics into SAM across all stages and a Fine-Tuning Module (FTM) that appends learnable tokens and applies a lightweight, shared MLP-based refinement to adapt SAM to human parsing. Together, these modules enable faster convergence and improved accuracy on Look into Person, Pascal-person-Part, and CIHP datasets, achieving state-of-the-art performance with notably reduced training time. The approach demonstrates the practical value of combining foundation models for domain-specific dense prediction tasks, offering a general recipe for efficient multi-modal adaptation in segmentation problems.

Abstract

Vision Foundation Model (VFM) such as the Segment Anything Model (SAM) and Contrastive Language-Image Pre-training Model (CLIP) has shown promising performance for segmentation and detection tasks. However, although SAM excels in fine-grained segmentation, it faces major challenges when applying it to semantic-aware segmentation. While CLIP exhibits a strong semantic understanding capability via aligning the global features of language and vision, it has deficiencies in fine-grained segmentation tasks. Human parsing requires to segment human bodies into constituent parts and involves both accurate fine-grained segmentation and high semantic understanding of each part. Based on traits of SAM and CLIP, we formulate high efficient modules to effectively integrate features of them to benefit human parsing. We propose a Semantic-Refinement Module to integrate semantic features of CLIP with SAM features to benefit parsing. Moreover, we formulate a high efficient Fine-tuning Module to adjust the pretrained SAM for human parsing that needs high semantic information and simultaneously demands spatial details, which significantly reduces the training time compared with full-time training and achieves notable performance. Extensive experiments demonstrate the effectiveness of our method on LIP, PPP, and CIHP databases.

SCHNet: SAM Marries CLIP for Human Parsing

TL;DR

Abstract

SCHNet: SAM Marries CLIP for Human Parsing

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (4)