Corruption Robust Offline Reinforcement Learning with Human Feedback

Debmalya Mandal; Andi Nika; Parameswaran Kamalaruban; Adish Singla; Goran Radanović

Corruption Robust Offline Reinforcement Learning with Human Feedback

Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban, Adish Singla, Goran Radanović

TL;DR

This work studies corruption-robust offline reinforcement learning from human feedback (RLHF) under a Huber contamination model where an $\varepsilon$-fraction of data may be corrupted. It develops a reduction-based framework that robustifies reward model learning via robust logistic regression, constructs a confidence set around the reward, and performs pessimistic offline RL over that set. The authors provide provable guarantees across three data-coverage regimes—Uniform Coverage, Low Relative Condition Number, and Bounded Generalized Coverage Ratio—achieving suboptimality that scales with $\varepsilon$ (up to log factors) and exploiting both zero-order and first-order offline RL oracles. By fusing PbRL with corruption-robust offline RL, the paper delivers principled, near-optimal strategies for RLHF under adversarial or noisy preferences. It also lays groundwork for future work on extending these results to general function approximation and trajectory-based rewards.

Abstract

We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedback about human preferences, an $\varepsilon$-fraction of the pairs is corrupted (e.g., feedback flipped or trajectory features manipulated), capturing an adversarial attack or noisy human preferences. We aim to design algorithms that identify a near-optimal policy from the corrupted data, with provable guarantees. Existing theoretical works have separately studied the settings of corruption robust RL (learning from scalar rewards directly under corruption) and offline RLHF (learning from human feedback without corruption); however, they are inapplicable to our problem of dealing with corrupted data in offline RLHF setting. To this end, we design novel corruption robust offline RLHF methods under various assumptions on the coverage of the data-generating distributions. At a high level, our methodology robustifies an offline RLHF framework by first learning a reward model along with confidence sets and then learning a pessimistic optimal policy over the confidence set. Our key insight is that learning optimal policy can be done by leveraging an offline corruption-robust RL oracle in different ways (e.g., zero-order oracle or first-order oracle), depending on the data coverage assumptions. To our knowledge, ours is the first work that provides provable corruption robust offline RLHF methods.

Corruption Robust Offline Reinforcement Learning with Human Feedback

TL;DR

This work studies corruption-robust offline reinforcement learning from human feedback (RLHF) under a Huber contamination model where an

-fraction of data may be corrupted. It develops a reduction-based framework that robustifies reward model learning via robust logistic regression, constructs a confidence set around the reward, and performs pessimistic offline RL over that set. The authors provide provable guarantees across three data-coverage regimes—Uniform Coverage, Low Relative Condition Number, and Bounded Generalized Coverage Ratio—achieving suboptimality that scales with

(up to log factors) and exploiting both zero-order and first-order offline RL oracles. By fusing PbRL with corruption-robust offline RL, the paper delivers principled, near-optimal strategies for RLHF under adversarial or noisy preferences. It also lays groundwork for future work on extending these results to general function approximation and trajectory-based rewards.

Abstract

-fraction of the pairs is corrupted (e.g., feedback flipped or trajectory features manipulated), capturing an adversarial attack or noisy human preferences. We aim to design algorithms that identify a near-optimal policy from the corrupted data, with provable guarantees. Existing theoretical works have separately studied the settings of corruption robust RL (learning from scalar rewards directly under corruption) and offline RLHF (learning from human feedback without corruption); however, they are inapplicable to our problem of dealing with corrupted data in offline RLHF setting. To this end, we design novel corruption robust offline RLHF methods under various assumptions on the coverage of the data-generating distributions. At a high level, our methodology robustifies an offline RLHF framework by first learning a reward model along with confidence sets and then learning a pessimistic optimal policy over the confidence set. Our key insight is that learning optimal policy can be done by leveraging an offline corruption-robust RL oracle in different ways (e.g., zero-order oracle or first-order oracle), depending on the data coverage assumptions. To our knowledge, ours is the first work that provides provable corruption robust offline RLHF methods.

Paper Structure (26 sections, 28 theorems, 170 equations, 1 table, 7 algorithms)

This paper contains 26 sections, 28 theorems, 170 equations, 1 table, 7 algorithms.

Introduction
Related Work
Preliminaries
Offline RLHF
Contamination Model
Parametric Markov Decision Processes
Robust RLHF with Uniform Coverage
Low Relative Condition Number
Bounded Generalized Coverage Ratio
Discussion and Future Work
Appendix
Missing Proofs from Section (3)
Convergence of Alternating Optimization
Proof of Lemma (3.2)
Proof of Theorem (3.3)
...and 11 more sections

Key Result

Lemma 3.2

Suppose assumption asn:uniform_coverage holds with $\xi \ge 5 \varepsilon$ and $N \ge \Omega\left(\frac{H^{3/2}}{\varepsilon^2}\left( d + \log(1/\delta) \right) \right)$. Then algorithm alg:alternating_optimization returns $\widehat{\theta}$, so that with probability at least $1-\delta$, we have

Theorems & Definitions (51)

Definition 2.3: Linear MDP jin2020provably
Lemma 3.2
Theorem 3.3
Proposition 3.5
Lemma 4.2
Theorem 4.3
Proposition 4.5
Theorem 5.1
Theorem 5.3
Proposition 5.4
...and 41 more

Corruption Robust Offline Reinforcement Learning with Human Feedback

TL;DR

Abstract

Corruption Robust Offline Reinforcement Learning with Human Feedback

Authors

TL;DR

Abstract

Table of Contents

Key Result

Theorems & Definitions (51)