Span-Based Optimal Sample Complexity for Average Reward MDPs

Matthew Zurek; Yudong Chen

Span-Based Optimal Sample Complexity for Average Reward MDPs

Matthew Zurek, Yudong Chen

TL;DR

This work establishes a complexity bound for the sample complexity of learning an $\varepsilon$-optimal policy in an average-reward Markov decision process (MDP) under a generative model and develops upper bounds on certain instance-dependent variance parameters in terms of the span parameter.

Abstract

We study the sample complexity of learning an $\varepsilon$-optimal policy in an average-reward Markov decision process (MDP) under a generative model. We establish the complexity bound $\widetilde{O}\left(SA\frac{H}{\varepsilon^2} \right)$, where $H$ is the span of the bias function of the optimal policy and $SA$ is the cardinality of the state-action space. Our result is the first that is minimax optimal (up to log factors) in all parameters $S,A,H$ and $\varepsilon$, improving on existing work that either assumes uniformly bounded mixing times for all policies or has suboptimal dependence on the parameters. Our result is based on reducing the average-reward MDP to a discounted MDP. To establish the optimality of this reduction, we develop improved bounds for $γ$-discounted MDPs, showing that $\widetilde{O}\left(SA\frac{H}{(1-γ)^2\varepsilon^2} \right)$ samples suffice to learn a $\varepsilon$-optimal policy in weakly communicating MDPs under the regime that $γ\geq 1 - \frac{1}{H}$, circumventing the well-known lower bound of $\widetildeΩ\left(SA\frac{1}{(1-γ)^3\varepsilon^2} \right)$ for general $γ$-discounted MDPs. Our analysis develops upper bounds on certain instance-dependent variance parameters in terms of the span parameter. These bounds are tighter than those based on the mixing time or diameter of the MDP and may be of broader use.

Span-Based Optimal Sample Complexity for Average Reward MDPs

TL;DR

This work establishes a complexity bound for the sample complexity of learning an

-optimal policy in an average-reward Markov decision process (MDP) under a generative model and develops upper bounds on certain instance-dependent variance parameters in terms of the span parameter.

Abstract

We study the sample complexity of learning an

-optimal policy in an average-reward Markov decision process (MDP) under a generative model. We establish the complexity bound

, where

is the span of the bias function of the optimal policy and

is the cardinality of the state-action space. Our result is the first that is minimax optimal (up to log factors) in all parameters

and

, improving on existing work that either assumes uniformly bounded mixing times for all policies or has suboptimal dependence on the parameters. Our result is based on reducing the average-reward MDP to a discounted MDP. To establish the optimality of this reduction, we develop improved bounds for

-discounted MDPs, showing that

samples suffice to learn a

-optimal policy in weakly communicating MDPs under the regime that

, circumventing the well-known lower bound of

for general

-discounted MDPs. Our analysis develops upper bounds on certain instance-dependent variance parameters in terms of the span parameter. These bounds are tighter than those based on the mixing time or diameter of the MDP and may be of broader use.

Paper Structure (11 sections, 11 theorems, 65 equations, 1 table, 2 algorithms)

This paper contains 11 sections, 11 theorems, 65 equations, 1 table, 2 algorithms.

Introduction
Related Work
Our Approach
Problem Setup
Main Results
Outline of Analysis
Proofs
Technical Lemmas
Proofs of Theorem \ref{['thm:DMDP_bound']} and \ref{['thm:main_theorem']}
Proof of Proposition \ref{['thm:optimal_tmix_bound']}
Conclusion

Key Result

Theorem 1

Suppose $H \leq \frac{1}{1-\gamma}$ and $\varepsilon \leq H$. There exists a constant $C_2 > 0$ such that, for any $\delta \in (0,1)$, if $n \geq C_2\frac{H}{(1-\gamma)^2\varepsilon^2} \log \left( \frac{S A}{(1-\gamma)\delta \varepsilon}\right)$, then with probability at least $1-\delta$, the policy

Theorems & Definitions (20)

Theorem 1: Sample Complexity of Discounted MDP
Theorem 2: Sample Complexity of Average-reward MDP
Proposition 1
Lemma 3
Lemma 4
proof
Lemma 5
Lemma 6
proof
Lemma 7
...and 10 more

Span-Based Optimal Sample Complexity for Average Reward MDPs

TL;DR

Abstract

Span-Based Optimal Sample Complexity for Average Reward MDPs

Authors

TL;DR

Abstract

Table of Contents

Key Result

Theorems & Definitions (20)