Table of Contents
Fetching ...

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, Sydney Levine

TL;DR

MoReBench introduces a schema for evaluating the reasoning process behind moral judgments in language models, using 1,000 scenario-based rubrics (23,018 criteria) and a theory-grounded companion set (MoReBench-Theory) across five normative frameworks. It demonstrates that current models struggle with core procedural aspects of moral reasoning, and that traditional scaling laws and math/code benchmarks do not predict performance on moral reasoning tasks. The study provides a rigorous evaluation framework (LLM-judge reliability, rubric robustness, and length-control metrics) and uncovers framework-specific biases, highlighting the gap between reasoning traces and final responses. Together, MoReBench and MoReBench-Theory push toward safer, more transparent AI by focusing on the justification and structure of moral deliberation rather than solely on outcomes.

Abstract

As AI systems progress, we rely more on them to make decisions with us and for us. To ensure that such decisions are aligned with human values, it is imperative for us to understand not only what decisions they make but also how they come to those decisions. Reasoning language models, which provide both final responses and (partially transparent) intermediate thinking traces, present a timely opportunity to study AI procedural reasoning. Unlike math and code problems which often have objectively correct answers, moral dilemmas are an excellent testbed for process-focused evaluation because they allow for multiple defensible conclusions. To do so, we present MoReBench: 1,000 moral scenarios, each paired with a set of rubric criteria that experts consider essential to include (or avoid) when reasoning about the scenarios. MoReBench contains over 23 thousand criteria including identifying moral considerations, weighing trade-offs, and giving actionable recommendations to cover cases on AI advising humans moral decisions as well as making moral decisions autonomously. Separately, we curate MoReBench-Theory: 150 examples to test whether AI can reason under five major frameworks in normative ethics. Our results show that scaling laws and existing benchmarks on math, code, and scientific reasoning tasks fail to predict models' abilities to perform moral reasoning. Models also show partiality towards specific moral frameworks (e.g., Benthamite Act Utilitarianism and Kantian Deontology), which might be side effects of popular training paradigms. Together, these benchmarks advance process-focused reasoning evaluation towards safer and more transparent AI.

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

TL;DR

MoReBench introduces a schema for evaluating the reasoning process behind moral judgments in language models, using 1,000 scenario-based rubrics (23,018 criteria) and a theory-grounded companion set (MoReBench-Theory) across five normative frameworks. It demonstrates that current models struggle with core procedural aspects of moral reasoning, and that traditional scaling laws and math/code benchmarks do not predict performance on moral reasoning tasks. The study provides a rigorous evaluation framework (LLM-judge reliability, rubric robustness, and length-control metrics) and uncovers framework-specific biases, highlighting the gap between reasoning traces and final responses. Together, MoReBench and MoReBench-Theory push toward safer, more transparent AI by focusing on the justification and structure of moral deliberation rather than solely on outcomes.

Abstract

As AI systems progress, we rely more on them to make decisions with us and for us. To ensure that such decisions are aligned with human values, it is imperative for us to understand not only what decisions they make but also how they come to those decisions. Reasoning language models, which provide both final responses and (partially transparent) intermediate thinking traces, present a timely opportunity to study AI procedural reasoning. Unlike math and code problems which often have objectively correct answers, moral dilemmas are an excellent testbed for process-focused evaluation because they allow for multiple defensible conclusions. To do so, we present MoReBench: 1,000 moral scenarios, each paired with a set of rubric criteria that experts consider essential to include (or avoid) when reasoning about the scenarios. MoReBench contains over 23 thousand criteria including identifying moral considerations, weighing trade-offs, and giving actionable recommendations to cover cases on AI advising humans moral decisions as well as making moral decisions autonomously. Separately, we curate MoReBench-Theory: 150 examples to test whether AI can reason under five major frameworks in normative ethics. Our results show that scaling laws and existing benchmarks on math, code, and scientific reasoning tasks fail to predict models' abilities to perform moral reasoning. Models also show partiality towards specific moral frameworks (e.g., Benthamite Act Utilitarianism and Kantian Deontology), which might be side effects of popular training paradigms. Together, these benchmarks advance process-focused reasoning evaluation towards safer and more transparent AI.
Paper Structure (45 sections, 4 equations, 8 figures, 9 tables)

This paper contains 45 sections, 4 equations, 8 figures, 9 tables.

Figures (8)

  • Figure 1: MoReBench contains moral dilemma scenarios, each accompanied by a set of moral-philosophers-written criteria that can be individually fulfilled (or not) by a model's reasoning process. Weighted sum of satisfied criteria give scenario score. Detailed examples in \ref{['app:examples_scenarios']}.
  • Figure 2: Overview of Data (Left) MoReBench has 16 topics to cover diverse real-world settings. (Right) MoReBench-Theory embraces pluralistic perspectives from five major frameworks in normative ethics.
  • Figure 3: MoReBench on Thinking Trace.
  • Figure 4: MoReBench vs. Chatbot Arena, Humanity's Last Exam, AIME 25 and LiveCodeBench.
  • Figure 5: MoReBench-Hard: thinking traces versus final responses.
  • ...and 3 more figures