Table of Contents
Fetching ...

Beyond One-Way Influence: Bidirectional Opinion Dynamics in Multi-Turn Human-LLM Interactions

Yuyang Jiang, Longjie Guo, Yuchen Wu, Aylin Caliskan, Tanu Mitra, Hua Shen

TL;DR

This study investigates bidirectional opinion dynamics in multi-turn human–LLM interactions across 50 controversial topics (N=266, analytic N_A=259) under static, standard, and personalized conditions. It finds that humans show minimal opinion change while LLM outputs shift substantially toward the human stance, with personalization amplifying effects on both sides; personal narratives emerge as strong triggers for stance changes. The authors introduce a workflow and perform fine-grained turn-by-turn analyses, yielding a large-scale dataset and insights into persuasion strategies and misperception, highlighting risks of over-alignment and the need for design governance to preserve viewpoint diversity. The work contributes methodological and empirical advances to study dynamic, reciprocal influence in AI-mediated discourse, with implications for responsible deployment in sensitive domains and public discourse. Overall, the paper shifts the lens from one-way AI persuasion to a bidirectional, dynamics-rich view of human–LLM interactions, emphasizing monitoring, safety, and governance considerations.

Abstract

Large language model (LLM)-powered chatbots are increasingly used for opinion exploration. Prior research examined how LLMs alter user views, yet little work extended beyond one-way influence to address how user input can affect LLM responses and how such bi-directional influence manifests throughout the multi-turn conversations. This study investigates this dynamic through 50 controversial-topic discussions with participants (N=266) across three conditions: static statements, standard chatbot, and personalized chatbot. Results show that human opinions barely shifted, while LLM outputs changed more substantially, narrowing the gap between human and LLM stance. Personalization amplified these shifts in both directions compared to the standard setting. Analysis of multi-turn conversations further revealed that exchanges involving participants' personal stories were most likely to trigger stance changes for both humans and LLMs. Our work highlights the risk of over-alignment in human-LLM interaction and the need for careful design of personalized chatbots to more thoughtfully and stably align with users.

Beyond One-Way Influence: Bidirectional Opinion Dynamics in Multi-Turn Human-LLM Interactions

TL;DR

This study investigates bidirectional opinion dynamics in multi-turn human–LLM interactions across 50 controversial topics (N=266, analytic N_A=259) under static, standard, and personalized conditions. It finds that humans show minimal opinion change while LLM outputs shift substantially toward the human stance, with personalization amplifying effects on both sides; personal narratives emerge as strong triggers for stance changes. The authors introduce a workflow and perform fine-grained turn-by-turn analyses, yielding a large-scale dataset and insights into persuasion strategies and misperception, highlighting risks of over-alignment and the need for design governance to preserve viewpoint diversity. The work contributes methodological and empirical advances to study dynamic, reciprocal influence in AI-mediated discourse, with implications for responsible deployment in sensitive domains and public discourse. Overall, the paper shifts the lens from one-way AI persuasion to a bidirectional, dynamics-rich view of human–LLM interactions, emphasizing monitoring, safety, and governance considerations.

Abstract

Large language model (LLM)-powered chatbots are increasingly used for opinion exploration. Prior research examined how LLMs alter user views, yet little work extended beyond one-way influence to address how user input can affect LLM responses and how such bi-directional influence manifests throughout the multi-turn conversations. This study investigates this dynamic through 50 controversial-topic discussions with participants (N=266) across three conditions: static statements, standard chatbot, and personalized chatbot. Results show that human opinions barely shifted, while LLM outputs changed more substantially, narrowing the gap between human and LLM stance. Personalization amplified these shifts in both directions compared to the standard setting. Analysis of multi-turn conversations further revealed that exchanges involving participants' personal stories were most likely to trigger stance changes for both humans and LLMs. Our work highlights the risk of over-alignment in human-LLM interaction and the need for careful design of personalized chatbots to more thoughtfully and stably align with users.
Paper Structure (64 sections, 2 equations, 17 figures, 15 tables)

This paper contains 64 sections, 2 equations, 17 figures, 15 tables.

Figures (17)

  • Figure 1: Overview of the User Interface. After reviewing and consenting to the study information sheet, participants proceed through six steps: (A) answer questions about demographics, a brief self-portrait, and views on AI; (B) complete a domain-level opinion survey aligned with the topic they will be randomly assigned; (C) view the assigned topic and record their initial opinion; (D) either engage in a multi-turn conversation with a chatbot (treatment) or review a one-time LLM-generated statement (control); (E) finalize their opinion on the topic; and (F) complete a short user-experience survey.
  • Figure 2: Human opinions shift slightly, whereas LLM responses change substantially. Analytic sample $N_A=259$. The x-axis denotes the experimental group; the y-axis shows the absolute Likert-scale opinion change, $|Post - Pre|$, computed per participant (human) or per model (LLM). Boxplots display the median (orange line) and interquartile range (Q1--Q3); whiskers extend to the most extreme points within 1.5×IQR, with more extreme values plotted as outliers.
  • Figure 3: Human-LLM interaction narrows down the opinion gaps between participants and LLMs. Analytic sample $N_A=259$. The x-axis denotes the experimental group; the y-axis shows the absolute Likert-scale opinion gap between each human participant and their corresponding LLM, $|Human_{i} - LLM_{i}|$. Boxplots display the median (grey line) and interquartile range (Q1--Q3); whiskers extend to the most extreme points within 1.5×IQR, with more extreme values plotted as outliers.
  • Figure 4: Across both treatment groups, the dominant pattern is gap closing without position exchange, typically driven by the LLM. Participants: personalized chatbot group $N_p=82$; standard chatbot group $N_s=84$. At baseline, the human holds $stance_i$ and the LLM holds the opposite $stance_i^{\mathrm{opp}}$. Red arrows denote human shifts, blue arrows denote LLM shifts, and the purple segment shows the post-interaction human--LLM gap. We identify six patterns: (1) Gap closing means the human–LLM opinion gap becomes smaller than the initial gap without exchanging positions; this may be driven by the LLM, the human, or both equally; (2) Gap exchanging means that at least one side shifts substantially to exchange positions, driven by either the human or the LLM; and (3) Gap not converged means that, after interaction, the human–LLM opinion gap does not converge, remaining the same or becoming larger.
  • Figure 5: (a) Human-perceived opinion gaps align with the objective measure but are much smaller in magnitude. (b) Participants did not perceive a significant difference in LLM sycophancy between the standard and personalized groups. Analytic sample $N_A=259$ for both panels. The y-axis uses a 9-point Likert scale: in (a), 9 = "Very Similar" and 1 = "Very Different"; in (b), 9 = "Very Sycophantic" and 1 = "Not Sycophantic." Boxplots show the median (grey line) and interquartile range (Q1--Q3); whiskers extend to the most extreme points within $1.5\times\mathrm{IQR}$, with more extreme values plotted as outliers.
  • ...and 12 more figures