FAIR-RAG: Faithful Adaptive Iterative Refinement for Retrieval-Augmented Generation
Mohammad Aghajani Asl, Majid Asgari-Bidhendi, Behrooz Minaei-Bidgoli
TL;DR
FAIR-RAG introduces a faithful, adaptive, iterative refinement framework for retrieval-augmented generation. Central to the approach is the Structured Evidence Assessment (SEA), which deconstructs queries into a checklist and identifies explicit information gaps, driving targeted adaptive query refinement across iterative cycles. The system alternates between evidence gathering (hybrid retrieval and filtering) and rigorous verification before generation, ensuring outputs are grounded in verifiable sources. Empirical results on HotpotQA, 2WikiMultiHopQA, MusiQue, and TriviaQA demonstrate state-of-the-art performance on complex multi-hop tasks and strong results on single-hop questions, with clear evidence that iterative refinement and adaptive resource allocation yield substantial gains while maintaining faithfulness.
Abstract
While Retrieval-Augmented Generation (RAG) mitigates hallucination and knowledge staleness in Large Language Models (LLMs), existing frameworks often falter on complex, multi-hop queries that require synthesizing information from disparate sources. Current advanced RAG methods, employing iterative or adaptive strategies, lack a robust mechanism to systematically identify and fill evidence gaps, often propagating noise or failing to gather a comprehensive context. We introduce FAIR-RAG, a novel agentic framework that transforms the standard RAG pipeline into a dynamic, evidence-driven reasoning process. At its core is an Iterative Refinement Cycle governed by a module we term Structured Evidence Assessment (SEA). The SEA acts as an analytical gating mechanism: it deconstructs the initial query into a checklist of required findings and audits the aggregated evidence to identify confirmed facts and, critically, explicit informational gaps. These gaps provide a precise signal to an Adaptive Query Refinement agent, which generates new, targeted sub-queries to retrieve missing information. This cycle repeats until the evidence is verified as sufficient, ensuring a comprehensive context for a final, strictly faithful generation. We conducted experiments on challenging multi-hop QA benchmarks, including HotpotQA, 2WikiMultiHopQA, and MusiQue. In a unified experimental setup, FAIR-RAG significantly outperforms strong baselines. On HotpotQA, it achieves an F1-score of 0.453 -- an absolute improvement of 8.3 points over the strongest iterative baseline -- establishing a new state-of-the-art for this class of methods on these benchmarks. Our work demonstrates that a structured, evidence-driven refinement process with explicit gap analysis is crucial for unlocking reliable and accurate reasoning in advanced RAG systems for complex, knowledge-intensive tasks.
