Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

Peng Gao; Yujian Lee; Xiaofeng Zhang; Zailong Chen; Hui Zhang

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

Peng Gao, Yujian Lee, Xiaofeng Zhang, Zailong Chen, Hui Zhang

TL;DR

The paper tackles the problem of long-range attention decay in RoPE-based LVLMs, which impairs cross-modal reasoning over distant token pairs. It introduces Three-step Decay-Resilience Strategies (T-DRS), an inference-only framework consisting of SD-DRS, DC-DRS, and reRD-DRS, to reinforce distant dependencies while preserving locality, forming $A^{\text{T-DRS}} = A + A^{sd} + A^{dc} + A^{re}$. Evaluated in a training-free setting on ScienceQA-IMG, GQA, TextVQA, and POPE, T-DRS yields consistent improvements across multiple LVLM backbones (e.g., LLaVA1.5-7B, InterVL2-8B, Qwen2.5-VL-7B), demonstrating generality and robustness in long-context VQA tasks. Overall, the approach enables more reliable global coherence in multimodal reasoning without model retraining, offering a practical plug-and-play enhancement for Vision-Language understanding and QA systems.

Abstract

Large Vision-Language Models (LVLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they still face critical challenges in modeling long-range dependencies under the usage of Rotary Positional Encoding (ROPE). Although it can facilitate precise modeling of token positions, it induces progressive attention decay as token distance increases, especially with progressive attention decay over distant token pairs, which severely impairs the model's ability to remember global context. To alleviate this issue, we propose inference-only Three-step Decay Resilience Strategies (T-DRS), comprising (1) Semantic-Driven DRS (SD-DRS), amplifying semantically meaningful but distant signals via content-aware residuals, (2) Distance-aware Control DRS (DC-DRS), which can purify attention by smoothly modulating weights based on positional distances, suppressing noise while preserving locality, and (3) re-Reinforce Distant DRS (reRD-DRS), consolidating the remaining informative remote dependencies to maintain global coherence. Together, the T-DRS recover suppressed long-range token pairs without harming local inductive biases. Extensive experiments on Vision Question Answering (VQA) benchmarks demonstrate that T-DRS can consistently improve performance in a training-free manner. The code can be accessed in https://github.com/labixiaoq-qq/Remember-me

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

TL;DR

Abstract

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (4)