Dynamically Augmented CVaR for MDPs

Eugene A. Feinberg; Rui Ding

Dynamically Augmented CVaR for MDPs

Eugene A. Feinberg, Rui Ding

TL;DR

The paper tackles risk-averse optimization in finite MDPs by introducing Dynamically Augmented CVaR (DCVaR), a time-consistent lower bound to static CVaR implemented via a Dynamically augmented Robust MDP (DRMDP). It defines DCVaR through a dynamic game against Nature, establishes a transformed DRMDP (DRMDP1) with concave-in-$y$ value functions, and presents Algorithm DCVaR that constructs nonrandomized policies minimizing DCVaR for finite and infinite horizons. The core theoretical contribution is linking DCVaR minimization to a mass-transfer problem, proving the algorithm’s correctness, and detailing the structural properties of the value functions and action sets that enable efficient computation. An extension to stochastic costs via state augmentation demonstrates the framework’s generality and applicability to more realistic risk-sensitive settings.

Abstract

This paper studies optimization of Conditional Value-at-Risk (CVaR) for Markov Decision Processes (MDPs) with finite state and action sets. It introduces the Dynamically augmented CVaR (DCVaR) risk measure and provides an algorithm for its optimization. This paper investigates a specially defined Robust MDP (RMDP), in which the state space is augmented with the tail risk level. This RMDP, which we call the Dynamically augmented RMDP (DRMDP), was introduced to the literature for calculations of optimal CVaR values by value iteration more than ten years ago, but, as was understood later, these value iterations compute lower bounds of minimal static CVaRs. DCVaR is defined as a time consistent version of the static CVaR, and it is a lower bound of the static CVaR. It also can be considered as a dynamic version of the nested CVaR. This paper provides an algorithm constructing a policy optimizing DCVaR of total discounted costs. The correctness of this algorithm is proved by studying a special mass transfer problem. The results on RMDPs needed for this paper are provided in the appendix.

Dynamically Augmented CVaR for MDPs

TL;DR

value functions, and presents Algorithm DCVaR that constructs nonrandomized policies minimizing DCVaR for finite and infinite horizons. The core theoretical contribution is linking DCVaR minimization to a mass-transfer problem, proving the algorithm’s correctness, and detailing the structural properties of the value functions and action sets that enable efficient computation. An extension to stochastic costs via state augmentation demonstrates the framework’s generality and applicability to more realistic risk-sensitive settings.

Abstract

Paper Structure (10 sections, 15 theorems, 100 equations, 1 figure)

This paper contains 10 sections, 15 theorems, 100 equations, 1 figure.

Introduction
Preliminaries: CVaR and MDPs
Static CVaR and Robust MDPs
Dynamically Augmented CVaR
Formulation of the Main Result
Properties of the Mass Transfer Problems Solved by Nature
Properties of Value Functions and Sets of Optimal Actions.
Proof of Theorem \ref{['TMAIN']}.
Extension to Stochastic Cost Functions
Results on Robust MDPs Used in this Paper

Key Result

Theorem 3.1

For every $N=1,2,\ldots$ or for $N=\infty,$ there exist a nonrandomized optimal policy $\phi\in\Pi$ for the CVaR optimization problem for which ${\rm CVaR}_\alpha(Z_N;P_x^\phi)={\rm CVaR}_\alpha(Z_N;x)$ for all $x\in\mathbb{X}.$ In addition,

Figures (1)

Figure 1: Demonstration of Cases of Algorithm CVaR for $t>0.$

Theorems & Definitions (29)

Theorem 3.1
proof
Theorem 3.2
proof
Theorem 3.3
proof
Corollary 3.4
Definition 4.1
Theorem 4.2
proof
...and 19 more

Dynamically Augmented CVaR for MDPs

TL;DR

Abstract

Dynamically Augmented CVaR for MDPs

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (1)

Theorems & Definitions (29)