CoRECT: A Framework for Evaluating Embedding Compression Techniques at Scale
L. Caspari, M. Dinzinger, K. Ghosh Dastidar, C. Fellicious, J. Mitrović, M. Granitzer
TL;DR
CoRECT addresses the challenge of evaluating embedding compression for dense retrieval by introducing a controlled, large-scale framework that jointly considers corpus complexity and model-specific behavior. It combines the CoRE dataset with BeIR, four embedding models, and eight compression methods to enable 40+ experimental configurations, using batchwise embedding generation, per-batch compression, and cosine retrieval evaluated with standard IR metrics. The study shows that corpus size and retrieval granularity significantly impact performance and that there is no universal best compression method; method effectiveness is highly model-dependent, though non-learned quantization often offers robust size reductions with minimal loss. The work provides practical guidance for choosing compression techniques in real-world, large-scale retrieval systems and makes the data and code openly available for reproducible, cross-model evaluation.
Abstract
Dense retrieval systems have proven to be effective across various benchmarks, but require substantial memory to store large search indices. Recent advances in embedding compression show that index sizes can be greatly reduced with minimal loss in ranking quality. However, existing studies often overlook the role of corpus complexity -- a critical factor, as recent work shows that both corpus size and document length strongly affect dense retrieval performance. In this paper, we introduce CoRECT (Controlled Retrieval Evaluation of Compression Techniques), a framework for large-scale evaluation of embedding compression methods, supported by a newly curated dataset collection. To demonstrate its utility, we benchmark eight representative types of compression methods. Notably, we show that non-learned compression achieves substantial index size reduction, even on up to 100M passages, with statistically insignificant performance loss. However, selecting the optimal compression method remains challenging, as performance varies across models. Such variability highlights the necessity of CoRECT to enable consistent comparison and informed selection of compression methods. All code, data, and results are available on GitHub and HuggingFace.
