Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
Bolei Ma, Yong Cao, Indira Sen, Anna-Carolina Haensch, Frauke Kreuter, Barbara Plank, Daniel Hershcovich
TL;DR
This position paper advocates embracing open-ended generation in LLM-based social simulations to better capture the diversity and reasoning of real populations. By connecting decades of social science open-ended survey practices with NLP capabilities, it argues that unconstrained text outputs enable richer measurement, exploration of minority views, and more realistic agent heterogeneity, while reducing researcher-imposed directive bias. The authors propose methodological frameworks that combine manual coding, qualitative analysis, and hybrid human–machine workflows, and outline concrete use-cases across populations, qualitative interviews, deliberation, and social-media simulations. They also discuss critical challenges—data, evaluation of infinite responses, and ethical risks—and call for benchmarks, multidimensional evaluation, bias-mitigation strategies, and transparent governance. Overall, the paper highlights practical avenues for integrating social science rigor with open-ended NLP to advance credible and versatile social simulations.
Abstract
Large Language Models (LLMs) are increasingly used to simulate public opinion and other social phenomena. Most current studies constrain these simulations to multiple-choice or short-answer formats for ease of scoring and comparison, but such closed designs overlook the inherently generative nature of LLMs. In this position paper, we argue that open-endedness, using free-form text that captures topics, viewpoints, and reasoning processes "in" LLMs, is essential for realistic social simulation. Drawing on decades of survey-methodology research and recent advances in NLP, we argue why this open-endedness is valuable in LLM social simulations, showing how it can improve measurement and design, support exploration of unanticipated views, and reduce researcher-imposed directive bias. It also captures expressiveness and individuality, aids in pretesting, and ultimately enhances methodological utility. We call for novel practices and evaluation frameworks that leverage rather than constrain the open-ended generative diversity of LLMs, creating synergies between NLP and social science.
