N-Gram Induction Heads for In-Context RL: Improving Stability and Reducing Data Needs

Ilya Zisman; Alexander Nikulin; Viacheslav Sinii; Denis Tarasov; Nikita Lyubaykin; Andrei Polubarov; Igor Kiselev; Vladislav Kurenkov

N-Gram Induction Heads for In-Context RL: Improving Stability and Reducing Data Needs

Ilya Zisman, Alexander Nikulin, Viacheslav Sinii, Denis Tarasov, Nikita Lyubaykin, Andrei Polubarov, Igor Kiselev, Vladislav Kurenkov

TL;DR

This paper tackles the data inefficiency and training instability of in-context reinforcement learning (ICRL) by integrating N-Gram Induction Heads (NGH) into transformer-based agents. Building on Algorithm Distillation, it adds an NGH layer that directly encodes n-gram patterns into attention, with n-gram matching extended to pixel-based observations via vector quantization. The results show substantial data efficiency gains (up to 27x less data in some tasks), faster hyperparameter optimization, and successful application to both discrete and image-based environments, while preserving or improving baseline performance. This method offers a practical, more scalable path to robust ICRL across diverse observation spaces and task distributions.

Abstract

In-context learning allows models like transformers to adapt to new tasks from a few examples without updating their weights, a desirable trait for reinforcement learning (RL). However, existing in-context RL methods, such as Algorithm Distillation (AD), demand large, carefully curated datasets and can be unstable and costly to train due to the transient nature of in-context learning abilities. In this work, we integrated the n-gram induction heads into transformers for in-context RL. By incorporating these n-gram attention patterns, we considerably reduced the amount of data required for generalization and eased the training process by making models less sensitive to hyperparameters. Our approach matches, and in some cases surpasses, the performance of AD in both grid-world and pixel-based environments, suggesting that n-gram induction heads could improve the efficiency of in-context RL.

N-Gram Induction Heads for In-Context RL: Improving Stability and Reducing Data Needs

TL;DR

Abstract

N-Gram Induction Heads for In-Context RL: Improving Stability and Reducing Data Needs

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (10)