Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach

Dengwang Tang; Rahul Jain; Botao Hao; Zheng Wen

Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach

Dengwang Tang, Rahul Jain, Botao Hao, Zheng Wen

TL;DR

This work tackles online reinforcement learning in infinite-horizon MDPs with an available offline dataset generated by an imperfect expert. It introduces an ideal Bayesian PSRL framework (inf-iPSRL) and a practical bootstrap-style variant (inf-iRLSVI) that leverage offline demonstrations through a competence-aware model of the expert. The authors derive a prior-dependent regret bound showing how offline data quality influences performance and show that strong demonstrations can yield near-constant regret as data grows. They further connect online RL with imitation learning, offering a principled way to fuse offline and online information and suggesting robust directions for universal RL that unify online, offline, and imitation paradigms. The results provide a quantitative understanding of data efficiency gains from offline sources and practical algorithms to realize them in infinite-horizon settings.

Abstract

In this paper, we study the problem of efficient online reinforcement learning in the infinite horizon setting when there is an offline dataset to start with. We assume that the offline dataset is generated by an expert but with unknown level of competence, i.e., it is not perfect and not necessarily using the optimal policy. We show that if the learning agent models the behavioral policy (parameterized by a competence parameter) used by the expert, it can do substantially better in terms of minimizing cumulative regret, than if it doesn't do that. We establish an upper bound on regret of the exact informed PSRL algorithm that scales as $\tilde{O}(\sqrt{T})$. This requires a novel prior-dependent regret analysis of Bayesian online learning algorithms for the infinite horizon setting. We then propose the Informed RLSVI algorithm to efficiently approximate the iPSRL algorithm.

Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach

TL;DR

Abstract

. This requires a novel prior-dependent regret analysis of Bayesian online learning algorithms for the infinite horizon setting. We then propose the Informed RLSVI algorithm to efficiently approximate the iPSRL algorithm.

Paper Structure (13 sections, 6 theorems, 66 equations, 2 algorithms)

This paper contains 13 sections, 6 theorems, 66 equations, 2 algorithms.

Introduction
Contributions.
Notations.
Preliminaries
The Infinite-horizon Informed PSRL Algorithm for Average MDPs
Prior Dependent Bound
Bounding the Estimation Error of Optimal Policy
An Approximate Bayesian Algorithm
The Informed RLSVI Algorithm for Average-reward MDPs
Informd RLSVI Bridges Online RL and Imitation Learning
Conclusions
Proof of Theorem \ref{['thm:regbound']}
Proof of Lemma \ref{['lem:piestimator']}

Key Result

Theorem 3.1

Let $K_T:=\max\{k:\sum_{l=1}^{k-1} T_k < T \}$. The Bayesian regret for inf-iPSRL algorithm satisfies where

Theorems & Definitions (14)

Remark 2.3
Theorem 3.1
Corollary 3.2
Remark 3.3
Corollary 3.4
Lemma 3.5
proof
Remark 3.6
Lemma 3.9
Remark 4.1
...and 4 more

Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach

TL;DR

Abstract

Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach

Authors

TL;DR

Abstract

Table of Contents

Key Result

Theorems & Definitions (14)