RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
Adina Yakefu, Bin Xie, Chongyang Xu, Enwen Zhang, Erjin Zhou, Fan Jia, Haitao Yang, Haoqiang Fan, Haowei Zhang, Hongyang Peng, Jing Tan, Junwen Huang, Kai Liu, Kaixin Liu, Kefan Gu, Qinglun Zhang, Ruitao Zhang, Saike Huang, Shen Cheng, Shuaicheng Liu, Tiancai Wang, Tiezhen Wang, Wei Sun, Wenbin Tang, Yajun Wei, Yang Chen, Youqiang Gui, Yucheng Zhao, Yunchao Ma, Yunfei Wei, Yunhuan Yang, Yutong Guo, Ze Chen, Zhengyuan Du, Ziheng Zhang, Ziming Liu, Ziwei Yan
TL;DR
RoboChallenge provides a real-robot, online evaluation framework to benchmark vision-language-action policies at scale on a fleet of 10 heterogeneous robots. It introduces a remote-robot paradigm with timestamped observations, asynchronous action queues, and accessible demonstration data to enable fair, reproducible, large-scale testing beyond simulators. The Table30 benchmark comprises 30 diverse tasks around a table, evaluated with a two-tier training setting (Task-specific and Generalist) and a progress-based grading scheme across 10 rollouts per task. Empirical results show strong performance by advanced baselines, particularly Pi0.5, while identifying persistent challenges in temporal reasoning and deformable object manipulation that motivate future improvements in realism-focused benchmarking and protocol design.
Abstract
Testing on real machines is indispensable for robotic control algorithms. In the context of learning-based algorithms, especially VLA models, demand for large-scale evaluation, i.e. testing a large number of models on a large number of tasks, is becoming increasingly urgent. However, doing this right is highly non-trivial, especially when scalability and reproducibility is taken into account. In this report, we describe our methodology for constructing RoboChallenge, an online evaluation system to test robotic control algorithms, and our survey of recent state-of-the-art VLA models using our initial benchmark Table30.
