LLM-augmented empirical game theoretic simulation for social-ecological systems
Jennifer Shi, Christopher K. Frantz, Christian Kimmich, Saba Siddiki, Atrisha Sarkar
TL;DR
The paper investigates integrating large language models (LLMs) with empirical game-theoretic analysis (EGTA) to simulate governance in social-ecological systems (SES). It compares four frameworks—procedural ABMs, generative ABMs, naive LLM-EGTA, and expert-guided LLM-EGTA—using a real-world Amu Darya case to assess sustainability, wealth distribution, and robustness under centralized vs. decentralized governance. Key contributions include two LLM-augmented EGTA pipelines (naive and expert-guided), evidence that framework choice critically shapes dynamics (e.g., potential collapse under naive LLM-EGTA vs. sustained mixed economies with expert guidance and Pigouvian taxes), and a robustness analysis showing that expert-in-the-loop guidance and internalized externalities are essential for plausible SES outcomes. The work highlights the practical importance of combining formal equilibrium analysis with human expertise when embedding LLMs into SES simulations for policy analysis and institutional design.
Abstract
Designing institutions for social-ecological systems requires models that capture heterogeneity, uncertainty, and strategic interaction. Multiple modeling approaches have emerged to meet this challenge, including empirical game-theoretic analysis (EGTA), which merges ABM's scale and diversity with game-theoretic models' formal equilibrium analysis. The newly popular class of LLM-driven simulations provides yet another approach, and it is not clear how these approaches can be integrated with one another, nor whether the resulting simulations produce a plausible range of behaviours for real-world social-ecological governance. To address this gap, we compare four LLM-augmented frameworks: procedural ABMs, generative ABMs, LLM-EGTA, and expert guided LLM-EGTA, and evaluate them on a real-world case study of irrigation and fishing in the Amu Darya basin under centralized and decentralized governance. Our results show: first, procedural ABMs, generative ABMs, and LLM-augmented EGTA models produce strikingly different patterns of collective behaviour, highlighting the value of methodological diversity. Second, inducing behaviour through system prompts in LLMs is less effective than shaping behaviour through parameterized payoffs in an expert-guided EGTA-based model.
