More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning

Kaiwen Wang; Owen Oertell; Alekh Agarwal; Nathan Kallus; Wen Sun

More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning

Kaiwen Wang, Owen Oertell, Alekh Agarwal, Nathan Kallus, Wen Sun

TL;DR

DistRL learns the full return distribution and the authors prove second-order, variance-dependent bounds for online and offline RL with function approximation, improving over prior small-loss guarantees. The framework extends to online RL on low-rank MDPs and offline RL under single-policy coverage, and introduces a novel first-/second-order gap-dependent bound for contextual bandits. A new variance-change-of-measure technique underpins the proofs, enabling bounds that scale with $\operatorname{Var}(Z)$ rather than $V^*$. Empirically, DistRL-based contextual bandits outperform squared-loss baselines, demonstrating practical impact for faster learning and robustness in real data settings.

Abstract

In this paper, we prove that Distributional Reinforcement Learning (DistRL), which learns the return distribution, can obtain second-order bounds in both online and offline RL in general settings with function approximation. Second-order bounds are instance-dependent bounds that scale with the variance of return, which we prove are tighter than the previously known small-loss bounds of distributional RL. To the best of our knowledge, our results are the first second-order bounds for low-rank MDPs and for offline RL. When specializing to contextual bandits (one-step RL problem), we show that a distributional learning based optimism algorithm achieves a second-order worst-case regret bound, and a second-order gap dependent bound, simultaneously. We also empirically demonstrate the benefit of DistRL in contextual bandits on real-world datasets. We highlight that our analysis with DistRL is relatively simple, follows the general framework of optimism in the face of uncertainty and does not require weighted regression. Our results suggest that DistRL is a promising framework for obtaining second-order bounds in general RL settings, thus further reinforcing the benefits of DistRL.

More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning

TL;DR

rather than

. Empirically, DistRL-based contextual bandits outperform squared-loss baselines, demonstrating practical impact for faster learning and robustness in real data settings.

Abstract

Paper Structure (37 sections, 25 theorems, 79 equations, 1 figure, 2 tables, 3 algorithms)

This paper contains 37 sections, 25 theorems, 79 equations, 1 figure, 2 tables, 3 algorithms.

Introduction
Related Works
Theory of DistRL.
Small-loss Bounds from DistRL.
Other second-order bounds.
Preliminaries
Contextual Bandits (CB).
Reinforcement Learning (RL).
Distributional RL.
Triangular Discrimination.
Warmup: Second-Order Bounds for CBs
Practical considerations.
Proof of \ref{['thm:second-order-cb']}
First and Second-Order Gap-Dependent Bounds
Second-Order Bounds for Online DistRL
...and 22 more sections

Key Result

Theorem 2.1

In online and offline RL, a second-order bound implies a first-order bound (with a worse universal constant). This is formalized in thm:second-order-implies-small-loss.

Figures (1)

Figure 1: Cost curves for the Housing task (lower is better).

Theorems & Definitions (44)

Theorem 2.1: Informal
Theorem 4.1
Lemma 4.1
Lemma 4.1
Theorem 4.2
Definition 5.2: $\ell_p$-distributional eluder dimension
Theorem 5.3: Second-order bounds for Online RL
Definition 5.4: Low-Rank MDP
Corollary 5.5: Second-Order PAC Bound for Low-Rank MDPs
Lemma 5.5: Change of Variance
...and 34 more

More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning

TL;DR

Abstract

More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (1)

Theorems & Definitions (44)