Optimal Attention Temperature Enhances In-Context Learning under Distribution Shift

Samet Demir; Zafer Dogan

Optimal Attention Temperature Enhances In-Context Learning under Distribution Shift

Samet Demir, Zafer Dogan

TL;DR

This work addresses how to enhance in-context learning robustness of pretrained Transformers when test data deviate from pretraining distributions. By adopting a linearized softmax attention model, it derives closed-form generalization-error expressions and proves that an optimal attention temperature $\tau$ exists to minimize error under distribution shift. The authors validate the theory with synthetic linear-regression tasks and large language models (GPT-2 and LLaMA-2-7B), showing that adjusting $\tau$ can substantially improve ICL performance in practical deployments. Overall, the paper provides a principled framework and actionable guidance for selecting attention temperature to bolster ICL under real-world distribution shifts.

Abstract

Pretrained Transformers excel at in-context learning (ICL), inferring new tasks from only a handful of examples. Yet, their ICL performance can degrade sharply under distribution shift between pretraining and test data, a regime increasingly common in real-world deployments. While recent empirical work hints that adjusting the attention temperature in the softmax can enhance Transformer performance, the attention temperature's role in ICL under distribution shift remains unexplored. This paper provides the first theoretical and empirical study of attention temperature for ICL under distribution shift. Using a simplified but expressive "linearized softmax" framework, we derive closed-form generalization error expressions and prove that shifts in input covariance or label noise substantially impair ICL, but that an optimal attention temperature exists which minimizes this error. We then validate our predictions through extensive simulations on linear regression tasks and large-scale experiments with GPT-2 and LLaMA2-7B on question-answering benchmarks. Our results establish attention temperature as a principled and powerful mechanism for improving the robustness of ICL in pretrained Transformers, advancing theoretical understanding and providing actionable guidance for selecting attention temperature in practice.

Optimal Attention Temperature Enhances In-Context Learning under Distribution Shift

TL;DR

Abstract

Optimal Attention Temperature Enhances In-Context Learning under Distribution Shift

TL;DR

Abstract

Paper Structure

Table of Contents

Key Result

Figures (6)

Theorems & Definitions (13)