Table of Contents
Fetching ...

Addressing the alignment problem in transportation policy making: an LLM approach

Xiaoyu Yan, Tianxing Dai, Yu Marco Nie

TL;DR

The paper investigates aligning transportation policy with public preferences by combining a utility-based travel demand model with a multi-agent referendum powered by large language models. Policies are defined by three levers $r,t,\tau$ at three levels, forming $K=27$ options, and decisions are aggregated via voting schemes including Ranked-Choice and approvals, across Chicago and Houston. Results show LLM agents generally align with model-based benchmarks but exhibit model-specific biases (GPT-4o being more decisive; Claude-3.5 more dispersed) and context sensitivity, with sentiment analyses offering explanatory insights. The study demonstrates the potential of LLM-driven participatory simulations for policy design while highlighting alignment challenges, framing effects, and the need for careful model and prompt design in real-world planning.

Abstract

A key challenge in transportation planning is that the collective preferences of heterogeneous travelers often diverge from the policies produced by model-driven decision tools. This misalignment frequently results in implementation delays or failures. Here, we investigate whether large language models (LLMs), noted for their capabilities in reasoning and simulating human decision-making, can help inform and address this alignment problem. We develop a multi-agent simulation in which LLMs, acting as agents representing residents from different communities in a city, participate in a referendum on a set of transit policy proposals. Using chain-of-thought reasoning, LLM agents provide ranked-choice or approval-based preferences, which are aggregated using instant-runoff voting (IRV) to model democratic consensus. We implement this simulation framework with both GPT-4o and Claude-3.5, and apply it for Chicago and Houston. Our findings suggest that LLM agents are capable of approximating plausible collective preferences and responding to local context, while also displaying model-specific behavioral biases and modest divergences from optimization-based benchmarks. These capabilities underscore both the promise and limitations of LLMs as tools for solving the alignment problem in transportation decision-making.

Addressing the alignment problem in transportation policy making: an LLM approach

TL;DR

The paper investigates aligning transportation policy with public preferences by combining a utility-based travel demand model with a multi-agent referendum powered by large language models. Policies are defined by three levers at three levels, forming options, and decisions are aggregated via voting schemes including Ranked-Choice and approvals, across Chicago and Houston. Results show LLM agents generally align with model-based benchmarks but exhibit model-specific biases (GPT-4o being more decisive; Claude-3.5 more dispersed) and context sensitivity, with sentiment analyses offering explanatory insights. The study demonstrates the potential of LLM-driven participatory simulations for policy design while highlighting alignment challenges, framing effects, and the need for careful model and prompt design in real-world planning.

Abstract

A key challenge in transportation planning is that the collective preferences of heterogeneous travelers often diverge from the policies produced by model-driven decision tools. This misalignment frequently results in implementation delays or failures. Here, we investigate whether large language models (LLMs), noted for their capabilities in reasoning and simulating human decision-making, can help inform and address this alignment problem. We develop a multi-agent simulation in which LLMs, acting as agents representing residents from different communities in a city, participate in a referendum on a set of transit policy proposals. Using chain-of-thought reasoning, LLM agents provide ranked-choice or approval-based preferences, which are aggregated using instant-runoff voting (IRV) to model democratic consensus. We implement this simulation framework with both GPT-4o and Claude-3.5, and apply it for Chicago and Houston. Our findings suggest that LLM agents are capable of approximating plausible collective preferences and responding to local context, while also displaying model-specific behavioral biases and modest divergences from optimization-based benchmarks. These capabilities underscore both the promise and limitations of LLMs as tools for solving the alignment problem in transportation decision-making.
Paper Structure (29 sections, 8 equations, 10 figures, 7 tables)

This paper contains 29 sections, 8 equations, 10 figures, 7 tables.

Figures (10)

  • Figure 1: Joint design of public transit service and policy.
  • Figure 2: Framework of the multi-agent LLM simulation
  • Figure 3: Comparison of $u_k$, $\gamma_k$, and $G_k$ against $U_k$ for different policy configurations in Chicago.
  • Figure 4: Voting counts of different voting types (single round).
  • Figure 5: Ranked-Choice voting outcomes generated by GPT-4o and Claude-3.5. Scenario CHI-com, ten rounds.
  • ...and 5 more figures