DiFiC: Your Diffusion Model Holds the Secret to Fine-Grained Clustering

Ruohong Yang; Peng Hu; Xi Peng; Xiting Liu; Yunfan Li

DiFiC: Your Diffusion Model Holds the Secret to Fine-Grained Clustering

Ruohong Yang, Peng Hu, Xi Peng, Xiting Liu, Yunfan Li

TL;DR

DiFiC tackles fine-grained clustering by leveraging a pre-trained text-to-image diffusion model to distill semantics from textual prompts rather than image features. The method combines semantic distillation, object-focused diffusion via attention masks, and neighborhood-guided clustering to produce compact, discriminative semantic representations $S^*$ that drive clustering. Empirical results on four fine-grained datasets show state-of-the-art ACC and NMI, with clear ablations validating the three-module design. This work highlights a novel use of diffusion models for discriminative tasks and suggests diffusion-based approaches can unlock finer semantic distinctions without heavy reliance on data augmentations.

Abstract

Fine-grained clustering is a practical yet challenging task, whose essence lies in capturing the subtle differences between instances of different classes. Such subtle differences can be easily disrupted by data augmentation or be overwhelmed by redundant information in data, leading to significant performance degradation for existing clustering methods. In this work, we introduce DiFiC a fine-grained clustering method building upon the conditional diffusion model. Distinct from existing works that focus on extracting discriminative features from images, DiFiC resorts to deducing the textual conditions used for image generation. To distill more precise and clustering-favorable object semantics, DiFiC further regularizes the diffusion target and guides the distillation process utilizing neighborhood similarity. Extensive experiments demonstrate that DiFiC outperforms both state-of-the-art discriminative and generative clustering methods on four fine-grained image clustering benchmarks. We hope the success of DiFiC will inspire future research to unlock the potential of diffusion models in tasks beyond generation. The code will be released.

DiFiC: Your Diffusion Model Holds the Secret to Fine-Grained Clustering

TL;DR

Abstract

DiFiC: Your Diffusion Model Holds the Secret to Fine-Grained Clustering

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (6)