Multilevel Picard scheme for solving high-dimensional drift control problems with state constraints
Yuan Zhong
TL;DR
This paper addresses high-dimensional drift control problems with state constraints enforced by reflections in the nonnegative orthant. It develops a nonparametric, simulation-based multilevel Picard (MLP) scheme that estimates the value function and its gradient through a fixed-point representation derived from a Feynman–Kac framework and a Bismut–Elworthy–Li identity for reflected diffusions. The key contributions include a contraction-based fixed-point characterization of the gradient, a rigorous MLP estimator with polynomial-in-dimension and inverse-accuracy complexity, and extensive numerical experiments up to dimension $20$ demonstrating practical feasibility and competitive performance. The results provide a scalable, nongrid-based approach for dynamic control in queueing networks and related systems, with potential extensions to time-discretized references and deep-learning hybrids. Overall, the work offers a rigorous, implementable method to overcome the curse of dimensionality in a class of constrained stochastic control problems, delivering actionable value-function and policy insights for high-dimensional settings.
Abstract
Motivated by applications to the dynamic control of queueing networks, we develop a simulation-based scheme, the so-called multilevel Picard (MLP) approximation, for solving high-dimensional drift control problems whose states are constrained to stay within the nonnegative orthant, over a finite time horizon. We prove that under suitable conditions, the MLP approximation overcomes the curse of dimensionality in the following sense: To approximate the value function and its gradient evaluated at a given time and state to within a prescribed accuracy $\varepsilon$, the computational complexity grows at most polynomially in the problem dimension $d$ and $1/\varepsilon$. To illustrate the effectiveness of the scheme, we carry out numerical experiments for a class of test problems that are related to the dynamic scheduling problem of parallel server systems in heavy traffic, and demonstrate that the scheme is computationally feasible up to dimension at least $20$.
