Table of Contents
Fetching ...

Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models

Benjamin Reichman, Adar Avsian, Larry Heck

TL;DR

This work uncovers that large language models internally organize emotion into a low-dimensional, directional subspace that remains stable across depth and generalizes across eight emotion datasets in five languages. It introduces ML-AURA, Centered-SVD, and Space Alignment to reveal a universal emotional manifold and demonstrates that a learned steering module can alter internal emotional perception while preserving semantics, with strong control over basic emotions across languages. The findings show that emotion is distributed across layers rather than localized, enabling high linear separability and robust cross-domain transfer. These results provide a structured, manipulable account of how LLMs internalize affect and offer avenues for interpretable and controllable affective reasoning in NLP systems.

Abstract

This work investigates how large language models (LLMs) internally represent emotion by analyzing the geometry of their hidden-state space. The paper identifies a low-dimensional emotional manifold and shows that emotional representations are directionally encoded, distributed across layers, and aligned with interpretable dimensions. These structures are stable across depth and generalize to eight real-world emotion datasets spanning five languages. Cross-domain alignment yields low error and strong linear probe performance, indicating a universal emotional subspace. Within this space, internal emotion perception can be steered while preserving semantics using a learned intervention module, with especially strong control for basic emotions across languages. These findings reveal a consistent and manipulable affective geometry in LLMs and offer insight into how they internalize and process emotion.

Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models

TL;DR

This work uncovers that large language models internally organize emotion into a low-dimensional, directional subspace that remains stable across depth and generalizes across eight emotion datasets in five languages. It introduces ML-AURA, Centered-SVD, and Space Alignment to reveal a universal emotional manifold and demonstrates that a learned steering module can alter internal emotional perception while preserving semantics, with strong control over basic emotions across languages. The findings show that emotion is distributed across layers rather than localized, enabling high linear separability and robust cross-domain transfer. These results provide a structured, manipulable account of how LLMs internalize affect and offer avenues for interpretable and controllable affective reasoning in NLP systems.

Abstract

This work investigates how large language models (LLMs) internally represent emotion by analyzing the geometry of their hidden-state space. The paper identifies a low-dimensional emotional manifold and shows that emotional representations are directionally encoded, distributed across layers, and aligned with interpretable dimensions. These structures are stable across depth and generalize to eight real-world emotion datasets spanning five languages. Cross-domain alignment yields low error and strong linear probe performance, indicating a universal emotional subspace. Within this space, internal emotion perception can be steered while preserving semantics using a learned intervention module, with especially strong control for basic emotions across languages. These findings reveal a consistent and manipulable affective geometry in LLMs and offer insight into how they internalize and process emotion.
Paper Structure (17 sections, 2 equations, 13 figures, 12 tables)

This paper contains 17 sections, 2 equations, 13 figures, 12 tables.

Figures (13)

  • Figure 1: Results of ML-AURA by layer and emotion. Results are in terms of percent of neurons with an AUROC score above $0.9$.
  • Figure 2: Emotion centroids plotted on the emotional axis found
  • Figure 3: Plot of projected emotions in t-SNE space.
  • Figure 4: Cosine similarity of emotional centroids between datasets for Llama3.1-Base.
  • Figure 5: Cosine similarity of emotional centroids between datasets for Llama3.1-Instruct.
  • ...and 8 more figures