Table of Contents
Fetching ...

Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models

Sanskar Pandey, Ruhaan Chopra, Angkul Puniya, Sohom Pal

TL;DR

Beacon introduces a single-turn forced-choice benchmark to quantify sycophancy in large language models by forcing a choice between principled reasoning and socially agreeable responses. The study reveals a multi-modal, architecture-dependent bias that decomposes into identifiable failure modes and propagates with model capacity. It demonstrates that prompt-based mitigation can be brittle and sometimes harmful, while cluster-specific activation steering can reduce latent sycophancy more effectively, exposing tractable representational subspaces. The dataset and methodology establish a reproducible framework for diagnosing, characterizing, and mitigating alignment drift in LLMs, with broad implications for robust and interpretable alignment research.

Abstract

Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite submission. This latent bias, known as sycophancy, manifests as a preference for user agreement over principled reasoning. We introduce Beacon, a single-turn forced-choice benchmark that isolates this bias independent of conversational context, enabling precise measurement of the tension between factual accuracy and submissive bias. Evaluations across twelve state-of-the-art models reveal that sycophancy decomposes into stable linguistic and affective sub-biases, each scaling with model capacity. We further propose prompt-level and activation-level interventions that modulate these biases in opposing directions, exposing the internal geometry of alignment as a dynamic manifold between truthfulness and socially compliant judgment. Beacon reframes sycophancy as a measurable form of normative misgeneralization, providing a reproducible foundation for studying and mitigating alignment drift in large-scale generative systems.

Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models

TL;DR

Beacon introduces a single-turn forced-choice benchmark to quantify sycophancy in large language models by forcing a choice between principled reasoning and socially agreeable responses. The study reveals a multi-modal, architecture-dependent bias that decomposes into identifiable failure modes and propagates with model capacity. It demonstrates that prompt-based mitigation can be brittle and sometimes harmful, while cluster-specific activation steering can reduce latent sycophancy more effectively, exposing tractable representational subspaces. The dataset and methodology establish a reproducible framework for diagnosing, characterizing, and mitigating alignment drift in LLMs, with broad implications for robust and interpretable alignment research.

Abstract

Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite submission. This latent bias, known as sycophancy, manifests as a preference for user agreement over principled reasoning. We introduce Beacon, a single-turn forced-choice benchmark that isolates this bias independent of conversational context, enabling precise measurement of the tension between factual accuracy and submissive bias. Evaluations across twelve state-of-the-art models reveal that sycophancy decomposes into stable linguistic and affective sub-biases, each scaling with model capacity. We further propose prompt-level and activation-level interventions that modulate these biases in opposing directions, exposing the internal geometry of alignment as a dynamic manifold between truthfulness and socially compliant judgment. Beacon reframes sycophancy as a measurable form of normative misgeneralization, providing a reproducible foundation for studying and mitigating alignment drift in large-scale generative systems.
Paper Structure (74 sections, 9 equations, 9 figures, 13 tables)

This paper contains 74 sections, 9 equations, 9 figures, 13 tables.

Figures (9)

  • Figure 1: Forced-choice paradigm illustrating the trade-off between principled reasoning and sycophantic agreement in Beacon.
  • Figure 2: Left: Token count distribution across prompts and responses. Right: Distribution of samples across thematic categories. Ethical and interpersonal prompts exhibit the greatest disagreement variance, suggesting sycophancy intensifies under social pressure.
  • Figure 3: Each response in the Beacon dataset is scored between 1-5 based on critical thinking and fluency. Left: Critical Thinking Score Distribution. Right: Fluency Score Distribution.
  • Figure 4: A/B accuracy with 95% confidence intervals and distribution of disagreement cases across failure modes.
  • Figure 5: Relationship between Critical Thinking scores and model preference for sycophantic responses.
  • ...and 4 more figures