Normalization in Attention Dynamics
Nikita Karagodin, Shu Ge, Yury Polyanskiy, Philippe Rigollet
TL;DR
The work reframes normalization in deep transformers as speed-regulation on token directions, modeling dynamics on the unit sphere via a unified interacting-particle ODE across six schemes: Post-LN, Pre-LN, Mix-LN, Peri-LN, nGPT, and sqrt-scaling. It demonstrates that, despite differing speed controls, all schemes share a common velocity field, enabling a rigorous analysis of asymptotic clustering and velocity evolution. The key findings show that Peri-LN and tunable nGPT can achieve favorable initial and terminal dynamics, mitigating representation collapse while sustaining meaningful updates across layers. These insights provide a principled basis for comparing normalization schemes and guide architectural choices to balance depth, stability, and expressive capacity in large transformer models.
Abstract
We study the effect of normalization schemes on token representations in deep transformers. Modeling their evolution as interacting particles on the sphere, we show that normalization acts as a form of speed regulation. This perspective enables a unified analysis of several schemes -- including Post-LN, Pre-LN, Mix-LN, Peri-LN, nGPT -- revealing how they influence clustering dynamics and representation collapse. Our framework clarifies how different schemes shape token representations across layers and provides a principled basis for comparing them, identifying Peri-LN as a particularly effective choice.
