Automated Social Science: Language Models as Scientist and Subjects

Benjamin S. Manning; Kehang Zhu; John J. Horton

Automated Social Science: Language Models as Scientist and Subjects

Benjamin S. Manning, Kehang Zhu, John J. Horton

TL;DR

This work presents an automated in silico social science framework that integrates structural causal models (SCMs) with large language models (LLMs) to generate hypotheses, construct interacting agent simulations, run experiments, and analyze results. It demonstrates that SCMs serve as a rigorous blueprint for experimental design and data analysis, enabling automated hypothesis testing across multiple social scenarios. While LLMs reliably predict the direction of effects, their magnitude predictions require conditioning on the fitted SCM, with theory-backed auction results aligning closely with simulated outcomes. The system showcases scalability, interactivity, and replicability, advancing AI-assisted social science toward continuous, prespecified experimentation and theory validation.

Abstract

We present an approach for automatically generating and testing, in silico, social scientific hypotheses. This automation is made possible by recent advances in large language models (LLM), but the key feature of the approach is the use of structural causal models. Structural causal models provide a language to state hypotheses, a blueprint for constructing LLM-based agents, an experimental design, and a plan for data analysis. The fitted structural causal model becomes an object available for prediction or the planning of follow-on experiments. We demonstrate the approach with several scenarios: a negotiation, a bail hearing, a job interview, and an auction. In each case, causal relationships are both proposed and tested by the system, finding evidence for some and not others. We provide evidence that the insights from these simulations of social interactions are not available to the LLM purely through direct elicitation. When given its proposed structural causal model for each scenario, the LLM is good at predicting the signs of estimated effects, but it cannot reliably predict the magnitudes of those estimates. In the auction experiment, the in silico simulation results closely match the predictions of auction theory, but elicited predictions of the clearing prices from the LLM are inaccurate. However, the LLM's predictions are dramatically improved if the model can condition on the fitted structural causal model. In short, the LLM knows more than it can (immediately) tell.

Automated Social Science: Language Models as Scientist and Subjects

TL;DR

Abstract

Paper Structure (34 sections, 3 equations, 19 figures, 3 tables)

This paper contains 34 sections, 3 equations, 19 figures, 3 tables.

Introduction
Overview of the system
Results of experiments
Bargaining over a mug
A bail hearing
Interviewing for a job as a lawyer
An auction for a piece of art
LLM predictions for paths and points
Predicting $y_i$
Predicting $\hat{\beta}$
Predicting $y_i|\hat{\beta}_{-i}$
Identifying causal structure ex-ante
Assuming causal structure from data
Searching for causal structure in data
Conclusion
...and 19 more sections

Figures (19)

Figure 1: An overview of the automated system.
Figure 2: Experimental design and fitted SCM for âtwo people bargaining over a mug.â
Figure 3: Experimental design and fitted SCM for âa judge is setting bail for a criminal defendant who committed 50,000 dollars in tax fraud.â
Figure 4: Experimental design and fitted SCM for âa person is interviewing for a job as a lawyer.â
Figure 5: Experimental design and fitted SCM for â3 bidders participating in an auction for a piece of art starting at fifty dollars.â
...and 14 more figures

Automated Social Science: Language Models as Scientist and Subjects

TL;DR

Abstract

Automated Social Science: Language Models as Scientist and Subjects

Authors

TL;DR

Abstract

Table of Contents

Figures (19)