Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models

Shuai Fu; Xiequn Wang; Qiushi Huang; Yu Zhang

Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models

Shuai Fu, Xiequn Wang, Qiushi Huang, Yu Zhang

TL;DR

This work reveals that soft-prompt vectors in vision-language models exhibit a Low-Norm Effect, where reducing norms at select prompt positions can boost performance while increasing norms can harm. To leverage this, it introduces Nemesis, a normalization framework with two losses (PUN and PAN) and a pre-inference step to identify normalization-worthy positions, enabling targeted regularization during soft-prompt tuning. Across 11 datasets and various few-shot and domain-generalization tasks, Nemesis consistently improves CoOp and extends to other PEFT methods like PLOT, underscoring the practical value of norm-aware soft-prompt tuning. The findings offer a principled direction for improving prompt-based adaptation of VLMs and invite further exploration of norm dynamics in soft-prompting.

Abstract

With the prevalence of large-scale pretrained vision-language models (VLMs), such as CLIP, soft-prompt tuning has become a popular method for adapting these models to various downstream tasks. However, few works delve into the inherent properties of learnable soft-prompt vectors, specifically the impact of their norms to the performance of VLMs. This motivates us to pose an unexplored research question: ``Do we need to normalize the soft prompts in VLMs?'' To fill this research gap, we first uncover a phenomenon, called the \textbf{Low-Norm Effect} by performing extensive corruption experiments, suggesting that reducing the norms of certain learned prompts occasionally enhances the performance of VLMs, while increasing them often degrades it. To harness this effect, we propose a novel method named \textbf{N}ormalizing th\textbf{e} soft-pro\textbf{m}pt v\textbf{e}ctors of vi\textbf{si}on-language model\textbf{s} (\textbf{Nemesis}) to normalize soft-prompt vectors in VLMs. To the best of our knowledge, our work is the first to systematically investigate the role of norms of soft-prompt vector in VLMs, offering valuable insights for future research in soft-prompt tuning. The code is available at \texttt{\href{https://github.com/ShyFoo/Nemesis}{https://github.com/ShyFoo/Nemesis}}.

Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models

TL;DR

Abstract

Paper Structure (29 sections, 5 equations, 6 figures, 15 tables, 1 algorithm)

This paper contains 29 sections, 5 equations, 6 figures, 15 tables, 1 algorithm.

Introduction
Low-Norm Effect
Methodology
A Revisit of Prompt-tuning Vision-language Models
Corruption Operations
The Nemesis Method
Experiments
Datasets
Implementation Details
Few-shot Image Recognition Results
Evaluation of Generalization Performance
In-depth Studies on the Low-Norm Effect in VLMs
Extendibility Analysis
Ablation studies and Hyper-parameter Analysis
Related Work
...and 14 more sections

Figures (6)

Figure 1: A schematic diagram of the Low-Norm Effect. (a) The occurrence of the Low-Norm Effect in soft-prompt tuning VLMs. (b) The occurrence frequency of the Low-Norm Effect across 11 datasets commonly used in soft-prompt tuning VLMs.
Figure 2: The few-shot recognition results of CoOp and CoOp+Nemesis (ours) on 11 datasets.
Figure 3: Analysis of the impact of soft prompt norms on the model's performance using the StanfordCars dataset as an example. The X-axis, Y1-axis, and Y2-axis represent training epochs, test accuracy, and average norm value of all soft-prompt vectors, respectively.
Figure A1: Comparison of the model's performance and norms of soft prompts before and after corrupting the prompt vector at various positions. The blue and green bars denote the model's performance and the norms of soft prompts, respectively. The lines represent the benchmarks of without corruption. Take the ImageNet dataset under 1 shot setting as an example.
Figure A2: The results of the Caltech101 dataset at the 1-shot setting during the training process.
...and 1 more figures

Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models

TL;DR

Abstract

Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models

Authors

TL;DR

Abstract

Table of Contents

Figures (6)