Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention

Yanbo Mao; Jianlong Fu; Ruoxuan Zhang; Hongxia Xie; Meibao Yao

Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention

Yanbo Mao, Jianlong Fu, Ruoxuan Zhang, Hongxia Xie, Meibao Yao

TL;DR

Imitation-learned robotic policies trained on mixed-quality demonstrations exhibit variable execution quality. The authors propose LIBERO-Elegant, an elegance-focused benchmark, and a decoupled refinement framework comprising an Elegance Critic trained via Calibrated Q-Learning and a Just-in-Time Intervention mechanism that selectively refines actions at decision-critical moments without retraining the base policy. Offline training on an Elegance-Enriched Dataset enables calibrated value estimation of action elegance, while JITI uses critic confidence to trigger high-cost, multi-sample refinement only when necessary. Results across simulation and real-world tasks show significant gains in Elegant Success Rate and strong generalization to unseen objects and contexts, highlighting the practical impact of optimizing not just task success but motion quality.

Abstract

Vision-Language-Action (VLA) models have enabled notable progress in general-purpose robotic manipulation, yet their learned policies often exhibit variable execution quality. We attribute this variability to the mixed-quality nature of human demonstrations, where the implicit principles that govern how actions should be carried out are only partially satisfied. To address this challenge, we introduce the LIBERO-Elegant benchmark with explicit criteria for evaluating execution quality. Using these criteria, we develop a decoupled refinement framework that improves execution quality without modifying or retraining the base VLA policy. We formalize Elegant Execution as the satisfaction of Implicit Task Constraints (ITCs) and train an Elegance Critic via offline Calibrated Q-Learning to estimate the expected quality of candidate actions. At inference time, a Just-in-Time Intervention (JITI) mechanism monitors critic confidence and intervenes only at decision-critical moments, providing selective, on-demand refinement. Experiments on LIBERO-Elegant and real-world manipulation tasks show that the learned Elegance Critic substantially improves execution quality, even on unseen tasks. The proposed model enables robotic control that values not only whether tasks succeed, but also how they are performed.

Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention

TL;DR

Abstract

Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (10)