Table of Contents
Fetching ...

Multilevel Picard scheme for solving high-dimensional drift control problems with state constraints

Yuan Zhong

TL;DR

This paper addresses high-dimensional drift control problems with state constraints enforced by reflections in the nonnegative orthant. It develops a nonparametric, simulation-based multilevel Picard (MLP) scheme that estimates the value function and its gradient through a fixed-point representation derived from a Feynman–Kac framework and a Bismut–Elworthy–Li identity for reflected diffusions. The key contributions include a contraction-based fixed-point characterization of the gradient, a rigorous MLP estimator with polynomial-in-dimension and inverse-accuracy complexity, and extensive numerical experiments up to dimension $20$ demonstrating practical feasibility and competitive performance. The results provide a scalable, nongrid-based approach for dynamic control in queueing networks and related systems, with potential extensions to time-discretized references and deep-learning hybrids. Overall, the work offers a rigorous, implementable method to overcome the curse of dimensionality in a class of constrained stochastic control problems, delivering actionable value-function and policy insights for high-dimensional settings.

Abstract

Motivated by applications to the dynamic control of queueing networks, we develop a simulation-based scheme, the so-called multilevel Picard (MLP) approximation, for solving high-dimensional drift control problems whose states are constrained to stay within the nonnegative orthant, over a finite time horizon. We prove that under suitable conditions, the MLP approximation overcomes the curse of dimensionality in the following sense: To approximate the value function and its gradient evaluated at a given time and state to within a prescribed accuracy $\varepsilon$, the computational complexity grows at most polynomially in the problem dimension $d$ and $1/\varepsilon$. To illustrate the effectiveness of the scheme, we carry out numerical experiments for a class of test problems that are related to the dynamic scheduling problem of parallel server systems in heavy traffic, and demonstrate that the scheme is computationally feasible up to dimension at least $20$.

Multilevel Picard scheme for solving high-dimensional drift control problems with state constraints

TL;DR

This paper addresses high-dimensional drift control problems with state constraints enforced by reflections in the nonnegative orthant. It develops a nonparametric, simulation-based multilevel Picard (MLP) scheme that estimates the value function and its gradient through a fixed-point representation derived from a Feynman–Kac framework and a Bismut–Elworthy–Li identity for reflected diffusions. The key contributions include a contraction-based fixed-point characterization of the gradient, a rigorous MLP estimator with polynomial-in-dimension and inverse-accuracy complexity, and extensive numerical experiments up to dimension demonstrating practical feasibility and competitive performance. The results provide a scalable, nongrid-based approach for dynamic control in queueing networks and related systems, with potential extensions to time-discretized references and deep-learning hybrids. Overall, the work offers a rigorous, implementable method to overcome the curse of dimensionality in a class of constrained stochastic control problems, delivering actionable value-function and policy insights for high-dimensional settings.

Abstract

Motivated by applications to the dynamic control of queueing networks, we develop a simulation-based scheme, the so-called multilevel Picard (MLP) approximation, for solving high-dimensional drift control problems whose states are constrained to stay within the nonnegative orthant, over a finite time horizon. We prove that under suitable conditions, the MLP approximation overcomes the curse of dimensionality in the following sense: To approximate the value function and its gradient evaluated at a given time and state to within a prescribed accuracy , the computational complexity grows at most polynomially in the problem dimension and . To illustrate the effectiveness of the scheme, we carry out numerical experiments for a class of test problems that are related to the dynamic scheduling problem of parallel server systems in heavy traffic, and demonstrate that the scheme is computationally feasible up to dimension at least .
Paper Structure (55 sections, 44 theorems, 416 equations, 3 figures, 4 tables)

This paper contains 55 sections, 44 theorems, 416 equations, 3 figures, 4 tables.

Key Result

Proposition 4.2

There exists a constant $C_\Gamma > 0$, which is independent of the problem dimension $d$, such that if $(h_i, g_i)$ is a solution to the Skorokhod problem for $f_i \in \mathbb{C}_+(\mathbb{R}^d)$, $i = 1,2$, then for all $t \in [0,\infty)$, $\| h_1 - h_2 \|_t + \| g_1 - g_2 \|_t \le C_\Gamma \| f_1

Figures (3)

  • Figure 1: Illustration of a $5$-dimensional open-chain system with $G$ defined in \ref{['eq:G']}. Server $i$ connects to buffers $i$ and $i+1$ for $i=1,\cdots,4$, and server $5$ connects to buffer $5$ only.
  • Figure 2: Policy plots under different $C_{\mathcal{A}}$ when $h_1 = 1.0$ and $h_2 = 2.0$. "+" indicates that for the estimated $\tilde{D}_1$ and $\tilde{D}_2$, $\tilde{D}_1- \tilde{D}_2 \leq 0$, and "x" indicates that $\tilde{D}_1 - \tilde{D}_2 > 0$. The policy interpretation is that at the "+" states, server $1$ helps to serve buffer $2$ at the maximum rate $C_{\mathcal{A}}$ possible, and does not do so at the "x" states.
  • Figure 3: Plot of value function differences between that under $C_{\mathcal{A}}=2.0$ and that under least control, when $h_1 = 1.0$ and $h_2 = 2.0$. Value function estimates for drift control with $C_{\mathcal{A}}=2.0$ are computed at level $n=4$. A negative value indicates that the value function estimate under drift control is lower than that under least control, and vice versa.

Theorems & Definitions (84)

  • Definition 4.1: Skorokhod Problem
  • Proposition 4.2: Lemma 3 of BlanchetChenSiGlynn2021
  • Definition 4.3
  • Proposition 4.4: Proposition 2.17 in LipshutzRamanan2018
  • Proposition 4.5: Proposition 6.3 in LipshutzRamanan2018
  • Definition 4.6
  • Lemma 4.7: Lemma 5.1 in LipshutzRamanan2018
  • Lemma 4.8: Lemma 5.2 in LipshutzRamanan2018
  • Proposition 4.9: Theorem 5.4 of LipshutzRamanan2018
  • Remark 4.10
  • ...and 74 more