RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling

Xiuying Wei; Anunay Yadav; Razvan Pascanu; Caglar Gulcehre

RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling

Xiuying Wei, Anunay Yadav, Razvan Pascanu, Caglar Gulcehre

TL;DR

RAT introduces a chunk-based temporal mixing layer that sits between RNNs and full self-attention, dividing sequences into chunks of length $L$ and applying intra-chunk recurrence alongside inter-chunk softmax attention to enable long-range retrieval with reduced computation. By tuning $L$, RAT interpolates between recurrence and attention, and a hybrid variant interleaves RAT with sliding-window attention to leverage strong local interactions. Extensive experiments on 1.3B-parameter models pretrained on 100B tokens show RAT with $L=16$ achieves substantial speedups (up to $\sim$7–10×) while maintaining competitive accuracy across short- and long-context benchmarks, and RAT-SWA often yields state-of-the-art results on long-context tasks. The work also provides thorough analyses of efficiency, ablations, and length-generalization strategies (RoPE/NoPE), highlighting RAT’s potential for scalable, efficient long-context language modeling and signaling avenues for future scaling and optimization.

Abstract

Transformers have become the cornerstone of modern large-scale language models, but their reliance on softmax attention poses a computational bottleneck at both training and inference. Recurrent models offer high efficiency, but compressing the full sequence into a fixed-size and holistic representation can suffer from memory degradation in long contexts and limit fine-grained retrieval. To address this, we propose RAT, an intermediate design that bridges the efficiency of RNNs and capacity of attention. RAT partitions the input into chunks, applies recurrence within each chunk for local dependencies, and softmax-based attention across chunks for long-range interactions. This design mitigates memory degradation and enables direct access to distant tokens, while retaining computational efficiency. Empirically, with a chunk size of 16, the RAT block achieves a 7$\times$ improvement in training speed for 100K sequence length and 9$times$ in generation at the 4K position, while maintaining similar performance compared to standard attention. We demonstrate this by training 1.3B parameter models from scratch and performing large-scale evaluations, including short- and long-context benchmarks, as well as supervised fine-tuning~(SFT). We further propose a hybrid architecture that interleaves RAT with local attention. By combining efficient long-range modeling with strong local interactions, this hybrid design not only improves inference speed and reduces cache memory usage, but also consistently enhances performance and shows the overall best results. Code is available at https://github.com/CLAIRE-Labo/RAT.

RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling

TL;DR

RAT introduces a chunk-based temporal mixing layer that sits between RNNs and full self-attention, dividing sequences into chunks of length

and applying intra-chunk recurrence alongside inter-chunk softmax attention to enable long-range retrieval with reduced computation. By tuning

, RAT interpolates between recurrence and attention, and a hybrid variant interleaves RAT with sliding-window attention to leverage strong local interactions. Extensive experiments on 1.3B-parameter models pretrained on 100B tokens show RAT with

achieves substantial speedups (up to

7–10×) while maintaining competitive accuracy across short- and long-context benchmarks, and RAT-SWA often yields state-of-the-art results on long-context tasks. The work also provides thorough analyses of efficiency, ablations, and length-generalization strategies (RoPE/NoPE), highlighting RAT’s potential for scalable, efficient long-context language modeling and signaling avenues for future scaling and optimization.

Abstract

improvement in training speed for 100K sequence length and 9

in generation at the 4K position, while maintaining similar performance compared to standard attention. We demonstrate this by training 1.3B parameter models from scratch and performing large-scale evaluations, including short- and long-context benchmarks, as well as supervised fine-tuning~(SFT). We further propose a hybrid architecture that interleaves RAT with local attention. By combining efficient long-range modeling with strong local interactions, this hybrid design not only improves inference speed and reduces cache memory usage, but also consistently enhances performance and shows the overall best results. Code is available at https://github.com/CLAIRE-Labo/RAT.

RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling

TL;DR

Abstract

RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (5)