WindTunnel -- A Framework for Community Aware Sampling of Large Corpora

Michael Iannelli

WindTunnel -- A Framework for Community Aware Sampling of Large Corpora

Michael Iannelli

TL;DR

WindTunnel is presented, a novel framework developed at Yext to generate representative samples of large corpora, enabling efficient end-to-end information retrieval experiments and overcomes limitations in current sampling methods, providing more accurate evaluations.

Abstract

Conducting comprehensive information retrieval experiments, such as in search or retrieval augmented generation, often comes with high computational costs. This is because evaluating a retrieval algorithm requires indexing the entire corpus, which is significantly larger than the set of (query, result) pairs under evaluation. This issue is especially pronounced in big data and neural retrieval, where indexing becomes increasingly time-consuming and complex. In this paper, we present WindTunnel, a novel framework developed at Yext to generate representative samples of large corpora, enabling efficient end-to-end information retrieval experiments. By preserving the community structure of the dataset, WindTunnel overcomes limitations in current sampling methods, providing more accurate evaluations.

WindTunnel -- A Framework for Community Aware Sampling of Large Corpora

TL;DR

Abstract

WindTunnel -- A Framework for Community Aware Sampling of Large Corpora

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (5)