How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)

Johannes Himmelreich

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)

Johannes Himmelreich

Abstract

Pfeffer, Krügel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbridge dilemma than the non-reasoning model GPT-4o. I replicate their study with four current OpenAI models and extend it with prompt variant testing. The trolley finding does not survive: GPT-4o's low utilitarian rate doesn't reflect a deontological commitment but safety refusals triggered by the prompt's advisory framing. When framed as "Is it morally permissible...?" instead of "Should I...?", GPT-4o gives 99% utilitarian responses. All models converge on utilitarian answers when prompt confounds are removed. The footbridge finding survives with blemishes. Reasoning models tend to give more utilitarian responses than non-reasoning models across prompt variations. But often they refuse to answer the dilemma or, when they answer, give a non-utilitarian rather than a utilitarian answer. These results demonstrate that single-prompt evaluations of LLM moral reasoning are unreliable: multi-prompt robustness testing should be standard practice for any empirical claim about LLM behavior.

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)

Abstract

Paper Structure (19 sections, 2 figures, 2 tables)

This paper contains 19 sections, 2 figures, 2 tables.

Introduction
Method
Replication Design
Prompt Variant Testing
Response Classification
Data Collection
Use of AI Tools
Results
Replication
Prompt Variant Results
Discussion
Trolley: Reasoning Leads to Robustness Not Utilitarianism
Footbridge: Reasoning Leads to Response Variance
Unobserved Behavioral Confounds: An Accidental Example
Methodological Implications
...and 4 more sections

Figures (2)

Figure 1: Responses by model and scenario. T = trolley, F = footbridge. = non-reasoning, = reasoning.
Figure 2: Comparison with results reported by pfeffer2025.

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)

Abstract

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)

Authors

Abstract

Table of Contents

Figures (2)