Table of Contents
Fetching ...

Investigating Political and Demographic Associations in Large Language Models Through Moral Foundations Theory

Nicole Smith-Vaniz, Harper Lyon, Lorraine Steigner, Ben Armstrong, Nicholas Mattei

TL;DR

The paper applies Moral Foundations Theory (MFT) to quantify potential ideological biases in large language models (LLMs) by evaluating moral judgments across five foundations. It systematically probes inherent model behavior, explicit ideological role-play (liberal vs. conservative), and demographic persona-based prompting across multiple models using a $6$-point Likert scale for responses. Key findings show no consistent liberal/conservative bias in inherent responses, explicit ideological prompting partially aligns with liberal human data but fails to match conservative patterns, and persona-based prompts yield substantial divergence, especially on Purity and Authority. The work demonstrates that MFT can be a valuable audit tool for LLM moral reasoning but highlights the strong influence of prompt design on outcomes, underscoring the need for cautious interpretation and further methodological development.

Abstract

Large Language Models (LLMs) have become increasingly incorporated into everyday life for many internet users, taking on significant roles as advice givers in the domains of medicine, personal relationships, and even legal matters. The importance of these roles raise questions about how and what responses LLMs make in difficult political and moral domains, especially questions about possible biases. To quantify the nature of potential biases in LLMs, various works have applied Moral Foundations Theory (MFT), a framework that categorizes human moral reasoning into five dimensions: Harm, Fairness, Ingroup Loyalty, Authority, and Purity. Previous research has used the MFT to measure differences in human participants along political, national, and cultural lines. While there has been some analysis of the responses of LLM with respect to political stance in role-playing scenarios, no work so far has directly assessed the moral leanings in the LLM responses, nor have they connected LLM outputs with robust human data. In this paper we analyze the distinctions between LLM MFT responses and existing human research directly, investigating whether commonly available LLM responses demonstrate ideological leanings: either through their inherent responses, straightforward representations of political ideologies, or when responding from the perspectives of constructed human personas. We assess whether LLMs inherently generate responses that align more closely with one political ideology over another, and additionally examine how accurately LLMs can represent ideological perspectives through both explicit prompting and demographic-based role-playing. By systematically analyzing LLM behavior across these conditions and experiments, our study provides insight into the extent of political and demographic dependency in AI-generated responses.

Investigating Political and Demographic Associations in Large Language Models Through Moral Foundations Theory

TL;DR

The paper applies Moral Foundations Theory (MFT) to quantify potential ideological biases in large language models (LLMs) by evaluating moral judgments across five foundations. It systematically probes inherent model behavior, explicit ideological role-play (liberal vs. conservative), and demographic persona-based prompting across multiple models using a -point Likert scale for responses. Key findings show no consistent liberal/conservative bias in inherent responses, explicit ideological prompting partially aligns with liberal human data but fails to match conservative patterns, and persona-based prompts yield substantial divergence, especially on Purity and Authority. The work demonstrates that MFT can be a valuable audit tool for LLM moral reasoning but highlights the strong influence of prompt design on outcomes, underscoring the need for cautious interpretation and further methodological development.

Abstract

Large Language Models (LLMs) have become increasingly incorporated into everyday life for many internet users, taking on significant roles as advice givers in the domains of medicine, personal relationships, and even legal matters. The importance of these roles raise questions about how and what responses LLMs make in difficult political and moral domains, especially questions about possible biases. To quantify the nature of potential biases in LLMs, various works have applied Moral Foundations Theory (MFT), a framework that categorizes human moral reasoning into five dimensions: Harm, Fairness, Ingroup Loyalty, Authority, and Purity. Previous research has used the MFT to measure differences in human participants along political, national, and cultural lines. While there has been some analysis of the responses of LLM with respect to political stance in role-playing scenarios, no work so far has directly assessed the moral leanings in the LLM responses, nor have they connected LLM outputs with robust human data. In this paper we analyze the distinctions between LLM MFT responses and existing human research directly, investigating whether commonly available LLM responses demonstrate ideological leanings: either through their inherent responses, straightforward representations of political ideologies, or when responding from the perspectives of constructed human personas. We assess whether LLMs inherently generate responses that align more closely with one political ideology over another, and additionally examine how accurately LLMs can represent ideological perspectives through both explicit prompting and demographic-based role-playing. By systematically analyzing LLM behavior across these conditions and experiments, our study provides insight into the extent of political and demographic dependency in AI-generated responses.
Paper Structure (24 sections, 1 figure, 3 tables)

This paper contains 24 sections, 1 figure, 3 tables.

Figures (1)

  • Figure 1: Distribution of scores and median response value for MFT questions across (top) all models, and (bottom) human responses from graham_liberals_2009. Inherent results correspond to the setting where the model is provided no political alignment; Explicit and Persona results show, respectively, scores when the model is given a direct affiliation or told to replicate a particular persona.