Practical Code RAG at Scale: Task-Aware Retrieval Design Choices under Compute Budgets
Timur Galimzyanov, Olga Kolomyttseva, Egor Bogomolov
TL;DR
This paper investigates how retrieval design choices affect code-focused RAG systems under realistic compute budgets, spanning two tasks from the Long Code Arena: PL→PL code completion and NL→PL bug localization. By systematically varying chunking strategies, similarity scorers, and tokenization granularity across context budgets from 128 to 16,384 tokens, the authors reveal task-dependent best practices: BM25 with word-level splitting delivers the best accuracy and efficiency for code-to-code retrieval, while dense embeddings provide superior cross-modal alignment for natural language-to-code tasks at higher latency. Chunk size should scale with available context, with 32–64 lines optimal for small budgets and whole-file retrieval competitive at 16K tokens; surprisingly, simple line-based chunking often matches or surpasses syntax-aware strategies. The study reports latency differences up to ~180× between configurations and offers evidence-based, task- and budget-aware guidelines for practical code RAG deployments, emphasizing the importance of balancing quality against compute cost in real-world settings.
Abstract
We study retrieval design for code-focused generation tasks under realistic compute budgets. Using two complementary tasks from Long Code Arena -- code completion and bug localization -- we systematically compare retrieval configurations across various context window sizes along three axes: (i) chunking strategy, (ii) similarity scoring, and (iii) splitting granularity. (1) For PL-PL, sparse BM25 with word-level splitting is the most effective and practical, significantly outperforming dense alternatives while being an order of magnitude faster. (2) For NL-PL, proprietary dense encoders (Voyager-3 family) consistently beat sparse retrievers, however requiring 100x larger latency. (3) Optimal chunk size scales with available context: 32-64 line chunks work best at small budgets, and whole-file retrieval becomes competitive at 16000 tokens. (4) Simple line-based chunking matches syntax-aware splitting across budgets. (5) Retrieval latency varies by up to 200x across configurations; BPE-based splitting is needlessly slow, and BM25 + word splitting offers the best quality-latency trade-off. Thus, we provide evidence-based recommendations for implementing effective code-oriented RAG systems based on task requirements, model constraints, and computational efficiency.
