TextCAM: Explaining Class Activation Map with Text

Qiming Zhao; Xingjian Li; Xiaoyu Cao; Xiaolong Wu; Min Xu

TextCAM: Explaining Class Activation Map with Text

Qiming Zhao, Xingjian Li, Xiaoyu Cao, Xiaolong Wu, Min Xu

TL;DR

TextCAM addresses the lack of semantic explanations in CAM-based visual explanations by fusing CAM with CLIP-based semantic space to produce textual rationales. It derives per-channel semantic vectors using CLIP embeddings and Linear Discriminant Analysis, and then combines them with CAM weights to generate $T_c(x)$; sparsity and decorrelation regularizers select diverse, concise phrases. The approach can group saliency into concept-based groups via a greedy channel assignment, producing multiple text-annotated saliency maps. Experiments on ImageNet, CLEVR, CUB, and DomainNet show TextCAM yields faithful, interpretable explanations, enables debiasing interventions, and transfers to Vision Transformers. The method is training-free and architecture-agnostic, offering a practical tool to diagnose and trust vision models.

Abstract

Deep neural networks (DNNs) have achieved remarkable success across domains but remain difficult to interpret, limiting their trustworthiness in high-stakes applications. This paper focuses on deep vision models, for which a dominant line of explainability methods are Class Activation Mapping (CAM) and its variants working by highlighting spatial regions that drive predictions. We figure out that CAM provides little semantic insight into what attributes underlie these activations. To address this limitation, we propose TextCAM, a novel explanation framework that enriches CAM with natural languages. TextCAM combines the precise spatial localization of CAM with the semantic alignment of vision-language models (VLMs). Specifically, we derive channel-level semantic representations using CLIP embeddings and linear discriminant analysis, and aggregate them with CAM weights to produce textual descriptions of salient visual evidence. This yields explanations that jointly specify where the model attends and what visual attributes likely support its decision. We further extend TextCAM to generate feature channels into semantically coherent groups, enabling more fine-grained visual-textual explanations. Experiments on ImageNet, CLEVR, and CUB demonstrate that TextCAM produces faithful and interpretable rationales that improve human understanding, detect spurious correlations, and preserve model fidelity.

TextCAM: Explaining Class Activation Map with Text

TL;DR

Abstract

TextCAM: Explaining Class Activation Map with Text

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (10)