Policy Iteration for Pareto-Optimal Policies in Stochastic Stackelberg Games

Mikoto Kudo; Yohei Akimoto

Policy Iteration for Pareto-Optimal Policies in Stochastic Stackelberg Games

Mikoto Kudo, Yohei Akimoto

TL;DR

The paper tackles policy design in general-sum stochastic Stackelberg games where SSEs may not exist, proposing Pareto-optimality as a robust alternative. It analyzes fixed-point operators, introduces state-dependent policies to achieve contraction and SSE-aligned fixed points, and develops a PO-focused policy iteration (PIPOP) with monotone improvement and convergence guarantees. The framework connects SSE existence to PO value functions and offers practical strategies (backtracking, scalarization) to ensure reasonable leader performance even when SSEs are absent. The work thus provides a theoretically grounded, practically applicable approach for leader-driven decision making under nonmyopic followers with known dynamics.

Abstract

In general-sum stochastic games, a stationary Stackelberg equilibrium (SSE) does not always exist, in which the leader maximizes leader's return for all the initial states when the follower takes the best response against the leader's policy. Existing methods of determining the SSEs require strong assumptions to guarantee the convergence and the coincidence of the limit with the SSE. Moreover, our analysis suggests that the performance at the fixed points of these methods is not reasonable when they are not SSEs. Herein, we introduced the concept of Pareto-optimality as a reasonable alternative to SSEs. We derive the policy improvement theorem for stochastic games with the best-response follower and propose an iterative algorithm to determine the Pareto-optimal policies based on it. Monotone improvement and convergence of the proposed approach are proved, and its convergence to SSEs is proved in a special case.

Policy Iteration for Pareto-Optimal Policies in Stochastic Stackelberg Games

TL;DR

Abstract

Paper Structure (35 sections, 10 theorems, 86 equations, 2 algorithms)

This paper contains 35 sections, 10 theorems, 86 equations, 2 algorithms.

Introduction
Preliminaries
Markov Decision Process
Stochastic Game and Stackelberg Game
Problem Setting
Notation
Related Works
Fixed Point Analysis of Operators
Operator defined by Bucarey2022-ro
State-dependent Markov policy and SSE Policy
Policy Iteration for Pareto-Optimal Policies
Pareto-Optimal Policy
Policy Improvement Theorem
Policy Iteration for PO Policies
Limitation and Future Work
...and 20 more sections

Key Result

Theorem 5.1

For $V:=(V_A,V_B)\in\mathcal{F}_{\mathcal{S}}\times\mathcal{F}_{\mathcal{S}}$, let $R_A(V)\in\mathcal{W}_A$ be a stationary policy whose action distribution on a state $s\in\mathcal{S}$ is in $R_A(s,V)$ and $R_B(f,V_B)$ be a deterministic stationary policy under $f\in\mathcal{W}_A$ whose action on a

Theorems & Definitions (26)

Definition 3.1: State value function
Definition 3.2: Follower's best response
Definition 3.3: SE value function
Definition 3.4: SSE policy
Theorem 5.1
Definition 5.2: (Leader's) State-dependent Markov policy
Theorem 5.3
Lemma 5.4
Theorem 5.5
Corollary 5.6
...and 16 more

Policy Iteration for Pareto-Optimal Policies in Stochastic Stackelberg Games

TL;DR

Abstract

Policy Iteration for Pareto-Optimal Policies in Stochastic Stackelberg Games

Authors

TL;DR

Abstract

Table of Contents

Key Result

Theorems & Definitions (26)