Monotonic Chunkwise Attention

Chung-Cheng Chiu; Colin Raffel

Monotonic Chunkwise Attention

Chung-Cheng Chiu, Colin Raffel

TL;DR

Problem: Standard soft attention incurs quadratic time/space and is unsuitable for online real-time transduction. Approach: MoChA combines a hard monotonic endpoint with soft attention over a small memory chunk preceding that endpoint, enabling online decoding with soft alignments and trainable via backpropagation. Findings: Achieves state-of-the-art performance on online speech recognition (WSJ) and substantially narrows the gap between monotonic and soft attention on document summarization (CNN/Daily Mail). Significance: Provides online, linear-time decoding while allowing local reorderings, with only modest additional parameters and computation.

Abstract

Sequence-to-sequence models with soft attention have been successfully applied to a wide variety of problems, but their decoding process incurs a quadratic time and space cost and is inapplicable to real-time sequence transduction. To address these issues, we propose Monotonic Chunkwise Attention (MoChA), which adaptively splits the input sequence into small chunks over which soft attention is computed. We show that models utilizing MoChA can be trained efficiently with standard backpropagation while allowing online and linear-time decoding at test time. When applied to online speech recognition, we obtain state-of-the-art results and match the performance of a model using an offline soft attention mechanism. In document summarization experiments where we do not expect monotonic alignments, we show significantly improved performance compared to a baseline monotonic attention-based model.

Monotonic Chunkwise Attention

TL;DR

Abstract

Monotonic Chunkwise Attention

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (4)