Positional Bias in Long-Document Ranking: Impact, Assessment, and Mitigation

Leonid Boytsov; David Akinpelu; Nipun Katyal; Tianyi Lin; Fangwei Gao; Yutian Zhao; Jeffrey Huang; Eric Nyberg

Positional Bias in Long-Document Ranking: Impact, Assessment, and Mitigation

Leonid Boytsov, David Akinpelu, Nipun Katyal, Tianyi Lin, Fangwei Gao, Yutian Zhao, Jeffrey Huang, Eric Nyberg

TL;DR

Positional bias in long-document ranking limits the apparent advantage of long-context transformers, motivating a benchmark-aware evaluation. The study systematically tests over 20 ranking methods across MS MARCO Documents, Robust04, BEIR, and a new MS MARCO FarRelevant diagnostic, finding that no long-context model beats the FirstP baseline by more than $5\%$ on average. It reveals that relevance tends to concentrate in the early document positions, a bias that persists across benchmarks and can cause overfitting to position; on FarRelevant, many long-context models perform at random without targeted debiasing or in-domain training. Debiasing training data yields mixed improvements, with PARADE and MaxP variants showing relative robustness, underscoring the need for careful benchmark design and more effective debiasing strategies. The work provides data and code to spur further research into robust long-context ranking and fair benchmarking.

Abstract

We tested over 20 Transformer models for ranking long documents (including recent LongP models trained with FlashAttention and RankGPT models "powered" by OpenAI and Anthropic cloud APIs). We compared them with the simple FirstP baseline, which applied the same model to truncated input (up to 512 tokens). On MS MARCO, TREC DL, and Robust04 no long-document model outperformed FirstP by more than 5% (on average). We hypothesized that this lack of improvement is not due to inherent model limitations, but due to benchmark positional bias (most relevant passages tend to occur early in documents), which is known to exist in MS MARCO. To confirm this, we analyzed positional relevance distributions across four long-document corpora (with six query sets) and observed the same early-position bias. Surprisingly, we also found bias in six BEIR collections, which are typically categorized as short-document datasets. We then introduced a new diagnostic dataset, MS MARCO FarRelevant, where relevant spans were deliberately placed beyond the first 512 tokens. On this dataset, many long-context models (including RankGPT) performed at random-baseline level, suggesting overfitting to positional bias. We also experimented with debiasing training data, but with limited success. Our findings (1) highlight the need for careful benchmark design in evaluating long-context models for document ranking, (2) identify model types that are more robust to positional bias, and (3) motivate further work on approaches to debias training data. We release our code and data to support further research.

Positional Bias in Long-Document Ranking: Impact, Assessment, and Mitigation

TL;DR

on average. It reveals that relevance tends to concentrate in the early document positions, a bias that persists across benchmarks and can cause overfitting to position; on FarRelevant, many long-context models perform at random without targeted debiasing or in-domain training. Debiasing training data yields mixed improvements, with PARADE and MaxP variants showing relative robustness, underscoring the need for careful benchmark design and more effective debiasing strategies. The work provides data and code to spur further research into robust long-context ranking and fair benchmarking.

Abstract

Paper Structure (37 sections, 1 equation, 8 figures, 11 tables, 1 algorithm)

This paper contains 37 sections, 1 equation, 8 figures, 11 tables, 1 algorithm.

Introduction
Related Work
Experiments
Data
Setup
Results
Realistic Data.
Synthetic Data.
Bias Mitigation.
Key Findings.
Conclusion
Limitations
Experimental Addendum: Training/Evaluation Setup, Ablations, and Detailed Results
Detailed Training and Evaluation Setup
General Setup
...and 22 more sections

Figures (8)

Figure 1: Positional relevance bias for three long-document collections (best viewed in color). We show a distribution of first relevant passage positions (red bars) vs. relevant document lengths (blue bars). Lengths and offsets are measured in the number of subword tokens (BERT-base tokenizer). See more results (including BEIR) in Figures \ref{['fig:relev_match_full_plot']} and \ref{['fig:relev_match_full_beir_plot']} in Appendix \ref{['sec:data_details']}.
Figure 2: Zero-shot vs. fine-tuned performance on MS MARCO FarRelevant. This figure shows results for a representative set of models.
Figure 3: Efficiency of long-document models vs respective (truncation) FirstP baselines. The figure shows an average relative gain (in %) vs. relative increase in run-time compared to respectiveFirstP baselines on MS MARCO, TREC DL 2019-2021, and Robust04 (for a representative subset of models). Except LongP RankGPT, LongP models truncate documents to be at most 1431 tokens. There is no truncation for RankGPT.
Figure 4: A sample relevant document for the Needle collection. The query/question is: "What is the Terracotta Army?". The answer-bearing sentence is marked by bold font.
Figure 5: A sample relevant document for the Passkey collection. The query/question is: "what is the passkey for Jimmy Moses?". The answer-bearing sentence is marked by bold font.
...and 3 more figures

Positional Bias in Long-Document Ranking: Impact, Assessment, and Mitigation

TL;DR

Abstract

Positional Bias in Long-Document Ranking: Impact, Assessment, and Mitigation

Authors

TL;DR

Abstract

Table of Contents

Figures (8)