Explain then Rank: Scale Calibration of Neural Rankers Using Natural Language Explanations from LLMs
Puxuan Yu, Daniel Cohen, Hemank Lamba, Joel Tetreault, Alex Jaimes
TL;DR
The paper tackles scale calibration for neural text rankers by leveraging natural language explanations (NLEs) generated by large language models (LLMs). It converts the ranking task to scoring NLEs via a two-part model: a frozen LLM g_{ Psi} that produces e^q for each query-document pair and a trainable ranker f_{\Theta} that scores the NLEs, i.e., $\phi_{\Phi}(q,\{d^q\}) = f_{\Theta}(g_{\Psi}(q,\{d^q\})) = f_{\Theta}(\{e^q\})$. The method explores literal and conditional prompting for NLEs, employs Monte Carlo sampling to form meta-NLEs, and aggregates multiple NLEs to encode uncertainty. Experiments on TREC and NTCIR show consistent improvements in calibration (CB‑ECE, ECE, MSE) and ranking (nDCG, nDCG@10), as well as downstream query performance prediction (QPP) metrics, relative to strong baselines, with reproducible results using Llama2-13B-Chat and BERT-based rankers. The approach offers a practical path to usable, interpretable ranking scores in large neural rankers, while acknowledging latency and bias considerations that invite future work on efficiency and robustness.
Abstract
In search settings, calibrating the scores during the ranking process to quantities such as click-through rates or relevance levels enhances a system's usefulness and trustworthiness for downstream users. While previous research has improved this notion of calibration for low complexity learning-to-rank models, the larger data demands and parameter count specific to modern neural text rankers produce unique obstacles that hamper the efficacy of methods intended for the learning-to-rank setting. This paper proposes exploiting large language models (LLMs) to provide relevance and uncertainty signals for these neural text rankers to produce scale-calibrated scores through Monte Carlo sampling of natural language explanations (NLEs). Our approach transforms the neural ranking task from ranking textual query-document pairs to ranking corresponding synthesized NLEs. Comprehensive experiments on two popular document ranking datasets show that the NLE-based calibration approach consistently outperforms past calibration methods and LLM-based methods for ranking, calibration, and query performance prediction tasks.
