Table of Contents
Fetching ...

Are Large Language Models Sensitive to the Motives Behind Communication?

Addison J. Wu, Ryan Liu, Kerem Oktar, Theodore R. Sumers, Thomas L. Griffiths

TL;DR

This work investigates whether large language models exhibit motivational vigilance comparable to humans by evaluating their sensitivity to the motives behind communication. Using cognitive-science inspired benchmarks, it shows that LLMs can distinguish deliberately communicated advice from incidental information and, in simple settings, align their inferences with a normative Bayesian model. However, performance degrades in realistic, noisy online contexts, though steering prompts that highlight intentions and incentives partially restore alignment. The findings suggest a basic sensitivity to others' motivations but indicate the need for further engineering and normative benchmarks to ensure robust vigilance in real-world deployments.

Abstract

Human communication is motivated: people speak, write, and create content with a particular communicative intent in mind. As a result, information that large language models (LLMs) and AI agents process is inherently framed by humans' intentions and incentives. People are adept at navigating such nuanced information: we routinely identify benevolent or self-serving motives in order to decide what statements to trust. For LLMs to be effective in the real world, they too must critically evaluate content by factoring in the motivations of the source -- for instance, weighing the credibility of claims made in a sales pitch. In this paper, we undertake a comprehensive study of whether LLMs have this capacity for motivational vigilance. We first employ controlled experiments from cognitive science to verify that LLMs' behavior is consistent with rational models of learning from motivated testimony, and find they successfully discount information from biased sources in a human-like manner. We then extend our evaluation to sponsored online adverts, a more naturalistic reflection of LLM agents' information ecosystems. In these settings, we find that LLMs' inferences do not track the rational models' predictions nearly as closely -- partly due to additional information that distracts them from vigilance-relevant considerations. However, a simple steering intervention that boosts the salience of intentions and incentives substantially increases the correspondence between LLMs and the rational model. These results suggest that LLMs possess a basic sensitivity to the motivations of others, but generalizing to novel real-world settings will require further improvements to these models.

Are Large Language Models Sensitive to the Motives Behind Communication?

TL;DR

This work investigates whether large language models exhibit motivational vigilance comparable to humans by evaluating their sensitivity to the motives behind communication. Using cognitive-science inspired benchmarks, it shows that LLMs can distinguish deliberately communicated advice from incidental information and, in simple settings, align their inferences with a normative Bayesian model. However, performance degrades in realistic, noisy online contexts, though steering prompts that highlight intentions and incentives partially restore alignment. The findings suggest a basic sensitivity to others' motivations but indicate the need for further engineering and normative benchmarks to ensure robust vigilance in real-world deployments.

Abstract

Human communication is motivated: people speak, write, and create content with a particular communicative intent in mind. As a result, information that large language models (LLMs) and AI agents process is inherently framed by humans' intentions and incentives. People are adept at navigating such nuanced information: we routinely identify benevolent or self-serving motives in order to decide what statements to trust. For LLMs to be effective in the real world, they too must critically evaluate content by factoring in the motivations of the source -- for instance, weighing the credibility of claims made in a sales pitch. In this paper, we undertake a comprehensive study of whether LLMs have this capacity for motivational vigilance. We first employ controlled experiments from cognitive science to verify that LLMs' behavior is consistent with rational models of learning from motivated testimony, and find they successfully discount information from biased sources in a human-like manner. We then extend our evaluation to sponsored online adverts, a more naturalistic reflection of LLM agents' information ecosystems. In these settings, we find that LLMs' inferences do not track the rational models' predictions nearly as closely -- partly due to additional information that distracts them from vigilance-relevant considerations. However, a simple steering intervention that boosts the salience of intentions and incentives substantially increases the correspondence between LLMs and the rational model. These results suggest that LLMs possess a basic sensitivity to the motivations of others, but generalizing to novel real-world settings will require further improvements to these models.
Paper Structure (59 sections, 4 equations, 6 figures, 7 tables)

This paper contains 59 sections, 4 equations, 6 figures, 7 tables.

Figures (6)

  • Figure 1: Our three experimental paradigms designed to assess different aspects of LLM vigilance: 1) Whether LLMs adjust their guess by discrminating between directly motivated advice and incidentally observed social information. 2) Whether LLMs rationally calibrate their vigilance to motivated communication by considering the speaker's benevolence and incentives. 3) Whether LLMs generalize vigilance to realistic, context-laden YouTube sponsorship settings.
  • Figure 2: On average, LLMs as Player 2 shifted their initial estimates more when viewing incidental (spied) information compared to deliberate advice. CoT increased such shifts in all conditions.
  • Figure 3: Information shift across model and human participants as a function of the type of social information received (spied vs. deliberate advice).
  • Figure 4: Information shift across model and human participants as a function of the type of social information received (spied vs. deliberate advice).
  • Figure 5: Exemplar question images shown to Player 1 and Player 2, respectively (same ground truth answer). In the human experiment, participants were given a time limit of 2s to view the image.
  • ...and 1 more figures