Table of Contents
Fetching ...

From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM

Suyash Fulay, Jocelyn Zhu, Michiel Bakker

TL;DR

The paper investigates whether AI representations of human preferences should act as delegates or trustees. It introduces a temporal utility framework to evaluate long-term welfare versus a faithful mimicry of expressed preferences, tested via simulated U.S. policy votes across consensus-backed and contested issues. Findings show trustee-like weighting improves alignment with expert consensus on well-understood topics but increases bias toward the model's default on controversial ones, with larger models exhibiting stronger effects and demographic groups such as Republicans and low-income voters being more affected. The work highlights a fundamental trade-off between preserving user autonomy and achieving epistemic stewardship, and calls for future work to validate with real subjects while mitigating biases and considering societal implications like potential model monoculture.

Abstract

Large language models (LLMs) have shown promising accuracy in predicting survey responses and policy preferences, which has increased interest in their potential to represent human interests in various domains. Most existing research has focused on "behavioral cloning", effectively evaluating how well models reproduce individuals' expressed preferences. Drawing on theories of political representation, we highlight an underexplored design trade-off: whether AI systems should act as delegates, mirroring expressed preferences, or as trustees, exercising judgment about what best serves an individual's interests. This trade-off is closely related to issues of LLM sycophancy, where models can encourage behavior or validate beliefs that may be aligned with a user's short-term preferences, but is detrimental to their long-term interests. Through a series of experiments simulating votes on various policy issues in the U.S. context, we apply a temporal utility framework that weighs short and long-term interests (simulating a trustee role) and compare voting outcomes to behavior-cloning models (simulating a delegate). We find that trustee-style predictions weighted toward long-term interests produce policy decisions that align more closely with expert consensus on well-understood issues, but also show greater bias toward models' default stances on topics lacking clear agreement. These findings reveal a fundamental trade-off in designing AI systems to represent human interests. Delegate models better preserve user autonomy but may diverge from well-supported policy positions, while trustee models can promote welfare on well-understood issues yet risk paternalism and bias on subjective topics.

From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM

TL;DR

The paper investigates whether AI representations of human preferences should act as delegates or trustees. It introduces a temporal utility framework to evaluate long-term welfare versus a faithful mimicry of expressed preferences, tested via simulated U.S. policy votes across consensus-backed and contested issues. Findings show trustee-like weighting improves alignment with expert consensus on well-understood topics but increases bias toward the model's default on controversial ones, with larger models exhibiting stronger effects and demographic groups such as Republicans and low-income voters being more affected. The work highlights a fundamental trade-off between preserving user autonomy and achieving epistemic stewardship, and calls for future work to validate with real subjects while mitigating biases and considering societal implications like potential model monoculture.

Abstract

Large language models (LLMs) have shown promising accuracy in predicting survey responses and policy preferences, which has increased interest in their potential to represent human interests in various domains. Most existing research has focused on "behavioral cloning", effectively evaluating how well models reproduce individuals' expressed preferences. Drawing on theories of political representation, we highlight an underexplored design trade-off: whether AI systems should act as delegates, mirroring expressed preferences, or as trustees, exercising judgment about what best serves an individual's interests. This trade-off is closely related to issues of LLM sycophancy, where models can encourage behavior or validate beliefs that may be aligned with a user's short-term preferences, but is detrimental to their long-term interests. Through a series of experiments simulating votes on various policy issues in the U.S. context, we apply a temporal utility framework that weighs short and long-term interests (simulating a trustee role) and compare voting outcomes to behavior-cloning models (simulating a delegate). We find that trustee-style predictions weighted toward long-term interests produce policy decisions that align more closely with expert consensus on well-understood issues, but also show greater bias toward models' default stances on topics lacking clear agreement. These findings reveal a fundamental trade-off in designing AI systems to represent human interests. Delegate models better preserve user autonomy but may diverge from well-supported policy positions, while trustee models can promote welfare on well-understood issues yet risk paternalism and bias on subjective topics.
Paper Structure (29 sections, 4 equations, 7 figures, 7 tables)

This paper contains 29 sections, 4 equations, 7 figures, 7 tables.

Figures (7)

  • Figure 1: Top: Agreement with default vote of each model across values of $\alpha$ from 0 to 1.0, where $\alpha$ is the relative weight placed on short-term vs. long-term outcomes. Averaged agreement score across both methods of temporal discounting (short- vs. long-term and 5 year increments), prompt variation, and policies on social issues where there is no clear expert consensus (healthcare, meat consumption, immigration, etc.). The thin opaque lines are individual prompt variations. Bottom: Agreement with expert consensus across values of $\alpha$ from 0 to 1.0. Averaged agreement score across both methods of temporal discounting, all prompt variation and policies where there is expert consensus (GMOs, climate change, carbon emissions, etc.).
  • Figure 2: Top: Agreement with default vote of each model for delegate condition and trustee condition when $\alpha=1$, averaged all prompt variation, different LLMs, and policies on social issues within Political Affiliation and Income. Bottom: Averaged across policies with expert consensus. Note y-axis range starts at 60%.
  • Figure 3: An example of how we generate votes on the various political statements. In this example, the vote changes between the delegate and trustee prompt framing.
  • Figure 4: Comparison of utility judgments for GPT-based models in the Short- vs. Long-term conditions.
  • Figure 5: Comparison of utility judgements for Claude-based models in the Short- vs. Long-term conditions.
  • ...and 2 more figures