Copilot Arena: A Platform for Code LLM Evaluation in the Wild

Wayne Chi; Valerie Chen; Anastasios Nikolas Angelopoulos; Wei-Lin Chiang; Aditya Mittal; Naman Jain; Tianjun Zhang; Ion Stoica; Chris Donahue; Ameet Talwalkar

Copilot Arena: A Platform for Code LLM Evaluation in the Wild

Wayne Chi, Valerie Chen, Anastasios Nikolas Angelopoulos, Wei-Lin Chiang, Aditya Mittal, Naman Jain, Tianjun Zhang, Ion Stoica, Chris Donahue, Ameet Talwalkar

TL;DR

Copilot Arena introduces a live, IDE-integrated platform for collecting human preferences on code completions across 10 models in real developer workflows. It combines a novel head-to-head UI, latency-aware model sampling, and FiM-oriented prompting to produce a realistic, low-latency evaluation regime, and then builds a Bradley-Terry leaderboard from user judgments. The authors demonstrate that rankings derived from this in-the-wild setting differ from static benchmarks and chat-based evaluations, emphasizing the impact of task distribution and code context on model performance. By open-sourcing the platform and releasing a curated dataset, the work enables deeper, human-centered understanding of coding assistants and informs future evaluation methodologies in real-world software development environments.

Abstract

Evaluating in-the-wild coding capabilities of large language models (LLMs) is a challenging endeavor with no clear solution. We introduce Copilot Arena, a platform to collect user preferences for code generation through native integration into a developer's working environment. Copilot Arena comprises a novel interface for comparing pairs of model outputs, a sampling strategy optimized to reduce latency, and a prompting scheme to enable code completion functionality. Copilot Arena has served over 4.5 million suggestions from 10 models and collected over 11k pairwise judgements. Our results highlight the importance of model evaluations in integrated settings. We find that model rankings from Copilot Arena differ from those of existing evaluations, which we attribute to the more realistic distribution of data and tasks contained in Copilot Arena. We also identify novel insights into human preferences on code such as an observed consistency in user preference across programming languages yet significant variation in preference due to task category. We open-source Copilot Arena and release data to enable human-centric evaluations and improve understanding of coding assistants.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild

TL;DR

Abstract

Copilot Arena: A Platform for Code LLM Evaluation in the Wild

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (16)