Reward Engineering for Spatial Epidemic Simulations: A Reinforcement Learning Platform for Individual Behavioral Learning

Radman Rakhshandehroo; Daniel Coombs

Reward Engineering for Spatial Epidemic Simulations: A Reinforcement Learning Platform for Individual Behavioral Learning

Radman Rakhshandehroo, Daniel Coombs

TL;DR

ContagionRL presents a Gymnasium-compatible platform to study reward engineering in spatial epidemic simulations by integrating a parameterizable SIRS+D model with reinforcement learning for a single learning agent in a population of non-learning humans. The methodology enables systematic evaluation of reward designs (e.g., constant, infection-probability-based, and a dense Potential Field reward) across multiple RL algorithms (PPO, SAC, A2C) and environmental conditions, including partial observability. Key findings show that the Potential Field reward supports superior policy learning through directional guidance and adherence incentives, while simpler rewards can lead to suboptimal or myopic strategies; partial observability can unexpectedly improve robustness and performance. The work contributes a modular, configurable platform for dissecting reward-behavior relationships in spatial epidemics, with implications for designing behaviorally informed interventions and understanding how information structure shapes adaptive responses.

Abstract

We present ContagionRL, a Gymnasium-compatible reinforcement learning platform specifically designed for systematic reward engineering in spatial epidemic simulations. Unlike traditional agent-based models that rely on fixed behavioral rules, our platform enables rigorous evaluation of how reward function design affects learned survival strategies across diverse epidemic scenarios. ContagionRL integrates a spatial SIRS+D epidemiological model with configurable environmental parameters, allowing researchers to stress-test reward functions under varying conditions including limited observability, different movement patterns, and heterogeneous population dynamics. We evaluate five distinct reward designs, ranging from sparse survival bonuses to a novel potential field approach, across multiple RL algorithms (PPO, SAC, A2C). Through systematic ablation studies, we identify that directional guidance and explicit adherence incentives are critical components for robust policy learning. Our comprehensive evaluation across varying infection rates, grid sizes, visibility constraints, and movement patterns reveals that reward function choice dramatically impacts agent behavior and survival outcomes. Agents trained with our potential field reward consistently achieve superior performance, learning maximal adherence to non-pharmaceutical interventions while developing sophisticated spatial avoidance strategies. The platform's modular design enables systematic exploration of reward-behavior relationships, addressing a knowledge gap in models of this type where reward engineering has received limited attention. ContagionRL is an effective platform for studying adaptive behavioral responses in epidemic contexts and highlight the importance of reward design, information structure, and environmental predictability in learning.

Reward Engineering for Spatial Epidemic Simulations: A Reinforcement Learning Platform for Individual Behavioral Learning

TL;DR

Abstract

Reward Engineering for Spatial Epidemic Simulations: A Reinforcement Learning Platform for Individual Behavioral Learning

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (15)