APCE: Adaptive Progressive Context Expansion for Long Context Processing
Baisub Lee, Sanghyun Byun, Mohanad Odema, Jung Guack, Jacob Song, Woo Seong Chung
TL;DR
The paper tackles the memory and performance degradation of long-context transformers (ContextRot) by introducing Adaptive Progressive Context Expansion (APCE), an input-chunk sparsification approach that selects the most semantically relevant chunks via a query-chunk similarity. APCE computes low-dimensional embeddings for input chunks and the current query, uses cosine similarity to pick the top-k chunks for attention, and supports reprioritization and asynchronous generation to adapt as decoding progresses, decoupling from hardware-specific kernels. This yields a drastic reduction in attention complexity from $O(N^2)$ to $O((k m)^2)$ and substantial memory savings (e.g., up to 55.6% in prefill and 32.8% in KV-cache) while matching or surpassing full-dense baselines on long-context summarization tasks like BookSum with context lengths up to 30k. The method is hardware-agnostic, complementary to existing sparsification techniques, and opens avenues for applying context-aware chunk selection to a broader set of long-context tasks, with tunable trade-offs via reprioritization and asynchronous generation.
Abstract
Deploying useful Long-Context Transformer Models (LCTMs) requires addressing two key challenges: (1) A growing memory footprint due to quadratic self-attention and linear KV-cache scaling in memory as sequence length increases; (2) the ContextRot phenomena where empirical evidence suggests that transformer architecture's performance degrades with increasing context length. Given the shared dependency on the input, a natural question arises: Can we surgically select the most important input chunks for processing to synergistically (a) reduce the memory footprint, and (b) mitigate the ContextRot effects? In this paper, we answer this question in the affirmative for long-context summarization tasks. We propose APCE as a context-aware solution to select the most important input chunks through low-dimensional semantic similarity matching with the current query. By directly operating on the input, APCE decouples from strict dependency on underlying hardware or CUDA environments, promising a compatible solution scalable to different deployment systems. Our empirical evaluations have demonstrated superior or on-par summarization performance for APCE compared to the full dense baseline using a fraction (50%-70%) of the input sequence resulting in KV-cache and self-attention memory efficiency improvements. We hope our findings inspire further research on context-aware efficiency solutions for LCTMs geared towards other relevant long-context tasks.
