Reinforcement Learning and Consumption-Savings Behavior
Brandon Kaplowitz
TL;DR
This paper investigates how reinforcement learning, implemented with a neural-network approximation of the continuation value, can explain two observed consumption patterns during downturns: higher MPCs among previously low-asset unemployed households and persistent consumption scarring from past unemployment. By modeling households as learning agents in a Markov decision process, the author shows that value-function approximation errors that evolve with experience can generate both higher MPCs and lower consumption, even without borrowing constraints. The approach reproduces key empirical findings from Ganong et al. (2024) and Malmendier & Shen (2024) within a unified, model-free learning framework, challenging traditional ex-ante heterogeneity or purely belief-updating explanations. The work highlights adaptive learning as a plausible mechanism shaping consumption beyond current income and assets, with implications for how policy transfers affect spending under uncertainty. Limitations include the quarterly frequency vs biennial data and the absence of explicit belief dynamics, suggesting avenues for richer model-based RL or POMDP formulations in future research.
Abstract
This paper demonstrates how reinforcement learning can explain two puzzling empirical patterns in household consumption behavior during economic downturns. I develop a model where agents use Q-learning with neural network approximation to make consumption-savings decisions under income uncertainty, departing from standard rational expectations assumptions. The model replicates two key findings from recent literature: (1) unemployed households with previously low liquid assets exhibit substantially higher marginal propensities to consume (MPCs) out of stimulus transfers compared to high-asset households (0.50 vs 0.34), even when neither group faces borrowing constraints, consistent with Ganong et al. (2024); and (2) households with more past unemployment experiences maintain persistently lower consumption levels after controlling for current economic conditions, a "scarring" effect documented by Malmendier and Shen (2024). Unlike existing explanations based on belief updating about income risk or ex-ante heterogeneity, the reinforcement learning mechanism generates both higher MPCs and lower consumption levels simultaneously through value function approximation errors that evolve with experience. Simulation results closely match the empirical estimates, suggesting that adaptive learning through reinforcement learning provides a unifying framework for understanding how past experiences shape current consumption behavior beyond what current economic conditions would predict.
