On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization

Ling Liang; Haizhao Yang

On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization

Ling Liang, Haizhao Yang

TL;DR

This work studies a general, regularized reward optimization problem in reinforcement learning, encompassing decision-dependent distributions. It analyzes a stochastic proximal gradient method and demonstrates an $O(\varepsilon^{-4})$ sample complexity to reach an $\varepsilon$-stationary point; it then introduces a variance-reduced PAGE-based estimator to reduce this to $O(\varepsilon^{-3})$ under additional assumptions, aligning with state-of-the-art results for discounted MDPs. The approach leverages Lipschitz smoothness of the objective $\mathcal{J}(\theta)$ and a proximal structure $\mathcal{G}(\theta)$, with an emphasis on non-oblivious RL settings and importance-weighted gradient estimation. The findings contribute a novel, theoretically-grounded variance-reduction pathway for general reward optimization in RL and offer insights into efficient sample use for proximal-gradient-based RL algorithms.

Abstract

We consider a regularized expected reward optimization problem in the non-oblivious setting that covers many existing problems in reinforcement learning (RL). In order to solve such an optimization problem, we apply and analyze the classical stochastic proximal gradient method. In particular, the method has shown to admit an $O(ε^{-4})$ sample complexity to an $ε$-stationary point, under standard conditions. Since the variance of the classical stochastic gradient estimator is typically large, which slows down the convergence, we also apply an efficient stochastic variance-reduce proximal gradient method with an importance sampling based ProbAbilistic Gradient Estimator (PAGE). Our analysis shows that the sample complexity can be improved from $O(ε^{-4})$ to $O(ε^{-3})$ under additional conditions. Our results on the stochastic (variance-reduced) proximal gradient method match the sample complexity of their most competitive counterparts for discounted Markov decision processes under similar settings. To the best of our knowledge, the proposed methods represent a novel approach in addressing the general regularized reward optimization problem.

On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization

TL;DR

sample complexity to reach an

-stationary point; it then introduces a variance-reduced PAGE-based estimator to reduce this to

under additional assumptions, aligning with state-of-the-art results for discounted MDPs. The approach leverages Lipschitz smoothness of the objective

and a proximal structure

, with an emphasis on non-oblivious RL settings and importance-weighted gradient estimation. The findings contribute a novel, theoretically-grounded variance-reduction pathway for general reward optimization in RL and offer insights into efficient sample use for proximal-gradient-based RL algorithms.

Abstract

sample complexity to an

-stationary point, under standard conditions. Since the variance of the classical stochastic gradient estimator is typically large, which slows down the convergence, we also apply an efficient stochastic variance-reduce proximal gradient method with an importance sampling based ProbAbilistic Gradient Estimator (PAGE). Our analysis shows that the sample complexity can be improved from

under additional conditions. Our results on the stochastic (variance-reduced) proximal gradient method match the sample complexity of their most competitive counterparts for discounted Markov decision processes under similar settings. To the best of our knowledge, the proposed methods represent a novel approach in addressing the general regularized reward optimization problem.

Paper Structure (14 sections, 7 theorems, 76 equations, 2 algorithms)

This paper contains 14 sections, 7 theorems, 76 equations, 2 algorithms.

Introduction
Related Work
Preliminary
The stochastic proximal gradient method
Variance reduction via PAGE
Conclusions
Proofs
Proof of Lemma \ref{['lemma-Lsmooth']}
Proof of Theorem \ref{['theorem-convergence']}
Proof of Lemma \ref{['lemma-bdd-var']}
Proof of Theorem \ref{['theorem-conv-pgd']}
Proof of Lemma \ref{['lemma-var-importance']}
Proof of Lemma \ref{['lemma-page-recursive']}
Proof Theorem \ref{['theorem-page-conv']}

Key Result

Lemma 4.1

Under Assumptions assumption-reward and assumption-policy, the gradient of $\mathcal{J}$ is $L$-smooth, i.e., with $L:= U(C_g^2+C_h) + \widetilde{C}_h + 2C_g\widetilde{C}_g > 0$.

Theorems & Definitions (21)

Example 3.3: MDP
Definition 3.4
Remark 3.5: Gradient mapping
Lemma 4.1
Remark 4.2: L-smoothness in MDPs
Theorem 4.3
Lemma 4.4
Theorem 4.5
Remark 4.6: Sample size
Remark 4.7: Global convergence
...and 11 more

On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization

TL;DR

Abstract

On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization

Authors

TL;DR

Abstract

Table of Contents

Key Result

Theorems & Definitions (21)