Hierarchical Prompt Decision Transformer: Improving Few-Shot Policy Generalization with Global and Adaptive Guidance

Zhe Wang; Haozhu Wang; Yanjun Qi

Hierarchical Prompt Decision Transformer: Improving Few-Shot Policy Generalization with Global and Adaptive Guidance

Zhe Wang, Haozhu Wang, Yanjun Qi

TL;DR

HPDT addresses the challenge of few-shot generalization in Offline Meta-RL by introducing hierarchical prompting for Decision Transformers. It learns a global token to capture task-level transition dynamics and rewards and retrieves adaptive tokens from demonstrations to provide timestep-specific guidance, aided by Time2Vec time embeddings. Across seven MuJoCo and MetaWorld tasks, HPDT consistently outperforms baselines, including PDT variants and full fine-tuning, demonstrating strong in-context learning with improved data efficiency. The approach offers a lightweight, context-aware mechanism for adapting policies to unseen tasks without requiring extensive retraining, advancing practical cross-task generalization in offline RL.

Abstract

Decision transformers recast reinforcement learning as a conditional sequence generation problem, offering a simple but effective alternative to traditional value or policy-based methods. A recent key development in this area is the integration of prompting in decision transformers to facilitate few-shot policy generalization. However, current methods mainly use static prompt segments to guide rollouts, limiting their ability to provide context-specific guidance. Addressing this, we introduce a hierarchical prompting approach enabled by retrieval augmentation. Our method learns two layers of soft tokens as guiding prompts: (1) global tokens encapsulating task-level information about trajectories, and (2) adaptive tokens that deliver focused, timestep-specific instructions. The adaptive tokens are dynamically retrieved from a curated set of demonstration segments, ensuring context-aware guidance. Experiments across seven benchmark tasks in the MuJoCo and MetaWorld environments demonstrate the proposed approach consistently outperforms all baseline methods, suggesting that hierarchical prompting for decision transformers is an effective strategy to enable few-shot policy generalization.

Hierarchical Prompt Decision Transformer: Improving Few-Shot Policy Generalization with Global and Adaptive Guidance

TL;DR

Abstract

Hierarchical Prompt Decision Transformer: Improving Few-Shot Policy Generalization with Global and Adaptive Guidance

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (4)