Fostering the Ecosystem of AI for Social Impact Requires Expanding and Strengthening Evaluation Standards
Bryan Wilder, Angela Zhou
TL;DR
The paper argues that impact in AI for social impact should extend beyond deploying novel methods and calls for expanding evaluation standards. It proposes three steps: recognize non-method contributions, recognize methodology-to-impact contributions even without deployment, and raise rigor for deployed evaluations. It details guidelines for pilot tests, randomized and non-randomized deployments, preregistration, power analysis, and event-study designs to guide both deployment and non-deployment research. By realigning incentives among authors, reviewers, and venues, the work aims to sustain a diverse and productive AISI research ecosystem that better serves partner organizations.
Abstract
There has been increasing research interest in AI/ML for social impact, and correspondingly more publication venues have refined review criteria for practice-driven AI/ML research. However, these review guidelines tend to most concretely recognize projects that simultaneously achieve deployment and novel ML methodological innovation. We argue that this introduces incentives for researchers that undermine the sustainability of a broader research ecosystem of social impact, which benefits from projects that make contributions on single front (applied or methodological) that may better meet project partner needs. Our position is that researchers and reviewers in machine learning for social impact must simultaneously adopt: 1) a more expansive conception of social impacts beyond deployment and 2) more rigorous evaluations of the impact of deployed systems.
