Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
Joongkyu Lee, Seouh-won Yi, Min-hwan Oh
TL;DR
This paper advances online preference-based reinforcement learning by modeling ranking feedback over action subsets via the Plackett–Luce model and proposing M-AUPO, which greedily selects assortments to maximize average uncertainty. The key theoretical contributions show that the suboptimality gap scales as $ ilde{O}ig(rac{d}{T}ig( extstyleig(rac{1}{|S_t|}ig)^{1/2}ig)ig)$, with an additional term that vanishes for large $T$, and crucially removes the customary $e^{B}$ dependence; a near-matching lower bound $oldsymbol{Omega}ig(rac{d}{K\, ext{sqrt}(T)}ig)$ confirms the benefit of larger subsets. The RB variant provides computational and empirical advantages, and a warm-up phase technique helps bound the leading term without auxiliary tricks. Experiments on synthetic data and real LLM-related datasets (TREC-DL and NECTAR) corroborate that increasing $K$ improves sample efficiency and that RB-based estimation yields strong practical performance. Overall, the work significantly clarifies how richer ranking feedback can tighten PbRL guarantees and informs scalable design for ranking-based learning-to-rank in contextual settings.
Abstract
We study online preference-based reinforcement learning (PbRL) with the goal of improving sample efficiency. While a growing body of theoretical work has emerged-motivated by PbRL's recent empirical success, particularly in aligning large language models (LLMs)-most existing studies focus only on pairwise comparisons. A few recent works (Zhu et al., 2023, Mukherjee et al., 2024, Thekumparampil et al., 2024) have explored using multiple comparisons and ranking feedback, but their performance guarantees fail to improve-and can even deteriorate-as the feedback length increases, despite the richer information available. To address this gap, we adopt the Plackett-Luce (PL) model for ranking feedback over action subsets and propose M-AUPO, an algorithm that selects multiple actions by maximizing the average uncertainty within the offered subset. We prove that M-AUPO achieves a suboptimality gap of $\tilde{O}\left( \frac{d}{T} \sqrt{ \sum_{t=1}^T \frac{1}{|S_t|}} \right)$, where $T$ is the total number of rounds, $d$ is the feature dimension, and $|S_t|$ is the size of the subset at round $t$. This result shows that larger subsets directly lead to improved performance and, notably, the bound avoids the exponential dependence on the unknown parameter's norm, which was a fundamental limitation in most previous works. Moreover, we establish a near-matching lower bound of $Ω\left( \frac{d}{K \sqrt{T}} \right)$, where $K$ is the maximum subset size. To the best of our knowledge, this is the first theoretical result in PbRL with ranking feedback that explicitly shows improved sample efficiency as a function of the subset size.
