Who's Gaming the System? A Causally-Motivated Approach for Detecting Strategic Adaptation
Trenton Chang, Lindsay Warrenburg, Sae-Hwan Park, Ravi B. Parikh, Maggie Makar, Jenna Wiens
TL;DR
Agents may manipulate inputs to ML-guided payouts, prompting a need to rank those most prone to gaming. The paper introduces a gaming deterrence parameter $\\lambda_p$ and shows that while $\\lambda_p$ is only partially identifiable, a ranking over agents is recoverable via counterfactual causal effects $\\tau(p,p')$. In synthetic experiments, causal-effect estimators outperform noncausal baselines in identifying top offenders; in a Medicare case study, the inferred rankings correlate with the prevalence of for-profit providers, suggesting real-world auditing utility. The approach relies on shared rewards, convex costs, and rational-actor assumptions, offering a principled framework for targeted audits while acknowledging potential limitations and ethical considerations.
Abstract
In many settings, machine learning models may be used to inform decisions that impact individuals or entities who interact with the model. Such entities, or agents, may game model decisions by manipulating their inputs to the model to obtain better outcomes and maximize some utility. We consider a multi-agent setting where the goal is to identify the "worst offenders:" agents that are gaming most aggressively. However, identifying such agents is difficult without knowledge of their utility function. Thus, we introduce a framework in which each agent's tendency to game is parameterized via a scalar. We show that this gaming parameter is only partially identifiable. By recasting the problem as a causal effect estimation problem where different agents represent different "treatments," we prove that a ranking of all agents by their gaming parameters is identifiable. We present empirical results in a synthetic data study validating the usage of causal effect estimation for gaming detection and show in a case study of diagnosis coding behavior in the U.S. that our approach highlights features associated with gaming.
