Table of Contents
Fetching ...

Sequentially Teaching Sequential Tasks $(ST)^2$: Teaching Robots Long-horizon Manipulation Skills

Zlatan Ajanović, Ravi Prakash, Leandro de Souza Rosa, Jens Kober

TL;DR

This work addresses the challenge of teaching long-horizon robotic manipulation where small deviations accumulate and teachers fatigue. It compares the traditional monolithic demonstration approach with a sequential segmentation framework, $ (ST)^2 $, allowing teachers to insert key-points and guide learning incrementally, with corrections localized to sub-tasks. A time-encoded Gaussian Process policy governs imitation learning, supported by a kinesthetic teaching interface and dynamic stiffness control that hands control back to humans when uncertainty rises. In a real-world supermarket restocking task with 16 participants, both methods achieve similar trajectory quality, but many users prefer $ (ST)^2 $ for its local correction opportunities and smoother learned behavior, suggesting that human-centric segmentation can enhance robustness and learning efficiency in long-horizon tasks.

Abstract

Learning from demonstration is effective for teaching robots complex skills with high sample efficiency. However, teaching long-horizon tasks with multiple skills is difficult, as deviations accumulate, distributional shift increases, and human teachers become fatigued, raising the chance of failure. In this work, we study user responses to two teaching frameworks: (i) a traditional monolithic approach, where users demonstrate the entire trajectory of a long-horizon task; and (ii) a sequential approach, where the task is segmented by the user and demonstrations are provided step by step. To support this study, we introduce $(ST)^2$, a sequential method for learning long-horizon manipulation tasks that allows users to control the teaching flow by defining key points, enabling incremental and structured demonstrations. We conducted a user study on a restocking task with 16 participants in a realistic retail environment to evaluate both user preference and method effectiveness. Our objective and subjective results show that both methods achieve similar trajectory quality and success rates. Some participants preferred the sequential approach for its iterative control, while others favored the monolithic approach for its simplicity.

Sequentially Teaching Sequential Tasks $(ST)^2$: Teaching Robots Long-horizon Manipulation Skills

TL;DR

This work addresses the challenge of teaching long-horizon robotic manipulation where small deviations accumulate and teachers fatigue. It compares the traditional monolithic demonstration approach with a sequential segmentation framework, , allowing teachers to insert key-points and guide learning incrementally, with corrections localized to sub-tasks. A time-encoded Gaussian Process policy governs imitation learning, supported by a kinesthetic teaching interface and dynamic stiffness control that hands control back to humans when uncertainty rises. In a real-world supermarket restocking task with 16 participants, both methods achieve similar trajectory quality, but many users prefer for its local correction opportunities and smoother learned behavior, suggesting that human-centric segmentation can enhance robustness and learning efficiency in long-horizon tasks.

Abstract

Learning from demonstration is effective for teaching robots complex skills with high sample efficiency. However, teaching long-horizon tasks with multiple skills is difficult, as deviations accumulate, distributional shift increases, and human teachers become fatigued, raising the chance of failure. In this work, we study user responses to two teaching frameworks: (i) a traditional monolithic approach, where users demonstrate the entire trajectory of a long-horizon task; and (ii) a sequential approach, where the task is segmented by the user and demonstrations are provided step by step. To support this study, we introduce , a sequential method for learning long-horizon manipulation tasks that allows users to control the teaching flow by defining key points, enabling incremental and structured demonstrations. We conducted a user study on a restocking task with 16 participants in a realistic retail environment to evaluate both user preference and method effectiveness. Our objective and subjective results show that both methods achieve similar trajectory quality and success rates. Some participants preferred the sequential approach for its iterative control, while others favored the monolithic approach for its simplicity.
Paper Structure (22 sections, 5 equations, 8 figures, 4 tables, 1 algorithm)

This paper contains 22 sections, 5 equations, 8 figures, 4 tables, 1 algorithm.

Figures (8)

  • Figure 1: Teaching a supermarket restocking task.
  • Figure 2: Restocking task. A) Initial state, milk carton in the box, B) Intermediate placement for re-grasping, C) Final state, milk carton on the shelf.
  • Figure 3: User $\#03$$(ST)^2$ teaching flow.
  • Figure 4: Distribution of users' preferences and subjective evaluation regarding the learning methods.
  • Figure 5: Quantitative workloads normalized (Z-score) per user. Smaller is better. Answers are divided according to the user's preferred method to highlight how the workloads affect their preferences. Users who preferred both methods are counted in both figures.
  • ...and 3 more figures