Influence Guided Context Selection for Effective Retrieval-Augmented Generation

Jiale Deng; Yanyan Shen; Ziyuan Pei; Youmin Chen; Linpeng Huang

Influence Guided Context Selection for Effective Retrieval-Augmented Generation

Jiale Deng, Yanyan Shen, Ziyuan Pei, Youmin Chen, Linpeng Huang

TL;DR

The paper tackles RAG hallucinations caused by noisy contexts by introducing Contextual Influence value (CI value), defined as φ_i(v) = v(C) − v(C ∖ {c_i}), which jointly captures query-, list-, and generator-aware signals. A hierarchical CI Surrogate Model (CSM) predicts CI values at inference time, with two training paradigms: supervised learning using oracle CI values and end-to-end training that leverages generator feedback via a differentiable masking mechanism. Experiments across 8 knowledge-intensive tasks and two LLM backbones show that CI-based context filtering substantially improves generation quality, achieving strong correlation to oracle CI (ρ > 0.75) and about a 15% average gain over baselines, while reducing latency. The work offers a scalable, hyperparameter-free approach to context selection in RAG, with potential for broad impact on grounding and reliability in knowledge-intensive NLP applications.

Abstract

Retrieval-Augmented Generation (RAG) addresses large language model (LLM) hallucinations by grounding responses in external knowledge, but its effectiveness is compromised by poor-quality retrieved contexts containing irrelevant or noisy information. While existing approaches attempt to improve performance through context selection based on predefined context quality assessment metrics, they show limited gains over standard RAG. We attribute this limitation to their failure in holistically utilizing available information (query, context list, and generator) for comprehensive quality assessment. Inspired by recent advances in data selection, we reconceptualize context quality assessment as an inference-time data valuation problem and introduce the Contextual Influence Value (CI value). This novel metric quantifies context quality by measuring the performance degradation when removing each context from the list, effectively integrating query-aware relevance, list-aware uniqueness, and generator-aware alignment. Moreover, CI value eliminates complex selection hyperparameter tuning by simply retaining contexts with positive CI values. To address practical challenges of label dependency and computational overhead, we develop a parameterized surrogate model for CI value prediction during inference. The model employs a hierarchical architecture that captures both local query-context relevance and global inter-context interactions, trained through oracle CI value supervision and end-to-end generator feedback. Extensive experiments across 8 NLP tasks and multiple LLMs demonstrate that our context selection method significantly outperforms state-of-the-art baselines, effectively filtering poor-quality contexts while preserving critical information. Code is available at https://github.com/SJTU-DMTai/RAG-CSM.

Influence Guided Context Selection for Effective Retrieval-Augmented Generation

TL;DR

Abstract

Influence Guided Context Selection for Effective Retrieval-Augmented Generation

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (12)

Theorems & Definitions (1)