Table of Contents
Fetching ...

Learning Mean-Field Games through Mean-Field Actor-Critic Flow

Mo Zhou, Haosheng Zhou, Ruimeng Hu

TL;DR

A rigorous convergence analysis using Lyapunov functionals is conducted and global exponential convergence of the MFAC flow is established under a suitable timescale, highlighting the algorithmic interplay among actor, critic, and distribution components.

Abstract

We propose the Mean-Field Actor-Critic (MFAC) flow, a continuous-time learning dynamics for solving mean-field games (MFGs), combining techniques from reinforcement learning and optimal transport. The MFAC framework jointly evolves the control (actor), value function (critic), and distribution components through coupled gradient-based updates governed by partial differential equations (PDEs). A central innovation is the Optimal Transport Geodesic Picard (OTGP) flow, which drives the distribution toward equilibrium along Wasserstein-2 geodesics. We conduct a rigorous convergence analysis using Lyapunov functionals and establish global exponential convergence of the MFAC flow under a suitable timescale. Our results highlight the algorithmic interplay among actor, critic, and distribution components. Numerical experiments illustrate the theoretical findings and demonstrate the effectiveness of the MFAC framework in computing MFG equilibria.

Learning Mean-Field Games through Mean-Field Actor-Critic Flow

TL;DR

A rigorous convergence analysis using Lyapunov functionals is conducted and global exponential convergence of the MFAC flow is established under a suitable timescale, highlighting the algorithmic interplay among actor, critic, and distribution components.

Abstract

We propose the Mean-Field Actor-Critic (MFAC) flow, a continuous-time learning dynamics for solving mean-field games (MFGs), combining techniques from reinforcement learning and optimal transport. The MFAC framework jointly evolves the control (actor), value function (critic), and distribution components through coupled gradient-based updates governed by partial differential equations (PDEs). A central innovation is the Optimal Transport Geodesic Picard (OTGP) flow, which drives the distribution toward equilibrium along Wasserstein-2 geodesics. We conduct a rigorous convergence analysis using Lyapunov functionals and establish global exponential convergence of the MFAC flow under a suitable timescale. Our results highlight the algorithmic interplay among actor, critic, and distribution components. Numerical experiments illustrate the theoretical findings and demonstrate the effectiveness of the MFAC framework in computing MFG equilibria.
Paper Structure (41 sections, 23 theorems, 361 equations, 13 figures, 1 algorithm)

This paper contains 41 sections, 23 theorems, 361 equations, 13 figures, 1 algorithm.

Key Result

Proposition 3.1

Under regularity conditions specified in Section sec:convergence, the derivative of $J^\mu[\alpha]$ with respect to $\alpha$ is

Figures (13)

  • Figure 1: Comparisons of value function gradients (top), equilibrium controls (bottom), and population measures in the systemic risk model (cf. Section \ref{['sec:SR']}). Five time snapshots are shown. Blue solid lines: baseline solutions; red solid lines: MFAC approximations; purple dashed lines: baseline densities; cyan histograms: empirical distributions from $5000$ sample paths of $\check{X}^m_t$.
  • Figure 2: Equilibrium population measures (left) and initial value functions (right) in the systemic risk model (cf. Section \ref{['sec:SR']}). Left: blue dashed lines denote baseline densities; red solid lines show kernel density estimations of $\tilde{\mu}_t$, computed from $5000$ LMC samples. Right: blue solid lines show the baseline value function; red solid lines show the MFAC approximation; purple dashed lines plot the initial density $\rho_0$.
  • Figure 3: Log-error curves in the systemic risk model (cf. Section \ref{['sec:SR']}) across different $\beta_\mu$. Errors are recorded every $10$ iterations.
  • Figure 4: Evolution of Lyapunov functions in the systemic risk model (cf. Section \ref{['sec:SR']}). Red: actor \ref{['eq:Lyapunov_actor']}; blue: critic \ref{['eq:Lyapunov_critic']}; orange: distribution term $\frac{1}{2}\mathcal{W}_2(\mu^\tau,\rho^{\mu^\tau,\alpha^\tau})$. Values are averaged over 10 independent runs and smoothed with a moving average (window size 10).
  • Figure 5: Comparisons of value function gradients (top), equilibrium controls (bottom), and population measures in the optimal execution problem (cf. Section \ref{['sec:Trader']}). Five time snapshots are shown. Blue solid lines: baseline solutions; red solid lines: MFAC approximations; purple dashed lines: baseline densities of control; cyan histograms: empirical distributions from 5000 sample paths of $\check{X}^m_t$.
  • ...and 8 more figures

Theorems & Definitions (46)

  • Definition 2.2: Mean-field equilibrium
  • Remark 2.3
  • Definition 2.4: Wasserstein-2 distance for measure flows
  • Proposition 3.1: Policy gradient theorem
  • Proposition 3.2
  • Remark 3.3
  • Theorem 4.4: Actor convergence
  • Theorem 4.5: Critic convergence
  • Theorem 4.6: Distribution convergence
  • Remark 4.7: OTGP for McKean--Vlasov SDEs
  • ...and 36 more