AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
Haoyu Huang, Hong Ting Tsang, Jiaxin Bai, Xi Peng, Gong Zhang, Yangqiu Song
TL;DR
The paper tackles the latency and retraining challenges of retrieval-augmented generation by proposing AtlasKV, a parametric approach that injects billion-scale knowledge graphs into LLMs. It introduces KG2KV to convert KG triples into Q-K-V data and HiKVP to prune KG keys hierarchically, enabling end-to-end integration under modest GPU memory. Empirical results show AtlasKV outperforms non-parametric and prior parametric baselines in knowledge grounding and generation relevance while dramatically reducing memory usage, even with 1B KG triples. This work enables scalable, training-free grounding of LLMs with massive KGs, facilitating efficient deployment in knowledge-intensive tasks without external retrievers.
Abstract
Retrieval-augmented generation (RAG) has shown some success in augmenting large language models (LLMs) with external knowledge. However, as a non-parametric knowledge integration paradigm for LLMs, RAG methods heavily rely on external retrieval modules and the retrieved textual context prior. Especially for very large scale knowledge augmentation, they would introduce substantial inference latency due to expensive searches and much longer relevant context. In this paper, we propose a parametric knowledge integration method, called \textbf{AtlasKV}, a scalable, effective, and general way to augment LLMs with billion-scale knowledge graphs (KGs) (e.g. 1B triples) using very little GPU memory cost (e.g. less than 20GB VRAM). In AtlasKV, we introduce KG2KV and HiKVP to integrate KG triples into LLMs at scale with sub-linear time and memory complexity. It maintains strong knowledge grounding and generalization performance using the LLMs' inherent attention mechanism, and requires no external retrievers, long context priors, or retraining when adapting to new knowledge.
