Paper Copilot: Tracking the Evolution of Peer Review in AI Conferences
Jing Yang, Qiyao Wei, Jiaxin Pei
TL;DR
Paper Copilot tackles the scalability and transparency of AI conference peer review by building a scalable, open archival system for reviews and an open dataset, complemented by longitudinal analyses of review dynamics (notably ICLR). It deploys a venue-configurable data-pipeline, a cross-venue, time-resolved dataset, and interactive analytics to study score dynamics, rebuttals, and talent trajectories. The work contributes a unified framework for tracking review evolution across venues and institutions, enabling reproducible meta-research and evidence-based improvements to the peer-review process, while addressing ethical, privacy, and bias considerations. Its findings reveal sharp, score-driven decision rules under high-volume pressure and distinct rebuttal dynamics, underscoring practical implications for fairness, accountability, and robustness in AI conference peer review.
Abstract
The rapid growth of AI conferences is straining an already fragile peer-review system, leading to heavy reviewer workloads, expertise mismatches, inconsistent evaluation standards, superficial or templated reviews, and limited accountability under compressed timelines. In response, conference organizers have introduced new policies and interventions to preserve review standards. Yet these ad-hoc changes often create further concerns and confusion about the review process, leaving how papers are ultimately accepted - and how practices evolve across years - largely opaque. We present Paper Copilot, a system that creates durable digital archives of peer reviews across a wide range of computer-science venues, an open dataset that enables researchers to study peer review at scale, and a large-scale empirical analysis of ICLR reviews spanning multiple years. By releasing both the infrastructure and the dataset, Paper Copilot supports reproducible research on the evolution of peer review. We hope these resources help the community track changes, diagnose failure modes, and inform evidence-based improvements toward a more robust, transparent, and reliable peer-review system.
