Sequentially Teaching Sequential Tasks $(ST)^2$: Teaching Robots Long-horizon Manipulation Skills
Zlatan Ajanović, Ravi Prakash, Leandro de Souza Rosa, Jens Kober
TL;DR
This work addresses the challenge of teaching long-horizon robotic manipulation where small deviations accumulate and teachers fatigue. It compares the traditional monolithic demonstration approach with a sequential segmentation framework, $ (ST)^2 $, allowing teachers to insert key-points and guide learning incrementally, with corrections localized to sub-tasks. A time-encoded Gaussian Process policy governs imitation learning, supported by a kinesthetic teaching interface and dynamic stiffness control that hands control back to humans when uncertainty rises. In a real-world supermarket restocking task with 16 participants, both methods achieve similar trajectory quality, but many users prefer $ (ST)^2 $ for its local correction opportunities and smoother learned behavior, suggesting that human-centric segmentation can enhance robustness and learning efficiency in long-horizon tasks.
Abstract
Learning from demonstration is effective for teaching robots complex skills with high sample efficiency. However, teaching long-horizon tasks with multiple skills is difficult, as deviations accumulate, distributional shift increases, and human teachers become fatigued, raising the chance of failure. In this work, we study user responses to two teaching frameworks: (i) a traditional monolithic approach, where users demonstrate the entire trajectory of a long-horizon task; and (ii) a sequential approach, where the task is segmented by the user and demonstrations are provided step by step. To support this study, we introduce $(ST)^2$, a sequential method for learning long-horizon manipulation tasks that allows users to control the teaching flow by defining key points, enabling incremental and structured demonstrations. We conducted a user study on a restocking task with 16 participants in a realistic retail environment to evaluate both user preference and method effectiveness. Our objective and subjective results show that both methods achieve similar trajectory quality and success rates. Some participants preferred the sequential approach for its iterative control, while others favored the monolithic approach for its simplicity.
