Table of Contents
Fetching ...

Learnable Mixed Nash Equilibria are Collectively Rational

Geelon So, Yi-An Ma

TL;DR

This work introduces uniform stability as a non-asymptotic stability notion for uncoupled learning dynamics in $N$-player normal-form games and links it to a form of collective rationality via strategic Pareto optimality. It shows that local uniform stability implies strategic Pareto optimality and equivalence to pointwise uniform stability, while non-uniform uniform stability precludes convergence to mixed equilibria; these insights are tied to the spectral properties of the game Jacobian. The authors analyze incremental smoothed best-response dynamics, proving global convergence to $eta$-smoothed equilibria under locally uniformly stable regions and establishing a $T^{-1/2}$ rate to Nash equilibria as $eta o 0$, with non-convergence results when uniform stability fails. They extend the theory to partially mixed equilibria via reduced games and linear steepness assumptions, and discuss implications, limitations, and directions for future work, including broader non-asymptotic analyses and relaxation of interaction assumptions. Overall, the paper argues that while dynamics around strict equilibria can yield socially inefficient outcomes, dynamics near mixed Nash equilibria exhibit collective rationality governed by uniform stability, providing a refined lens on learnability and social outcomes in uncoupled settings.

Abstract

We extend the study of learning in games to dynamics that exhibit non-asymptotic stability. We do so through the notion of uniform stability, which is concerned with equilibria of individually utility-seeking dynamics. Perhaps surprisingly, it turns out to be closely connected to economic properties of collective rationality. Under mild non-degeneracy conditions and up to strategic equivalence, if a mixed equilibrium is not uniformly stable, then it is not weakly Pareto optimal: there is a way for all players to improve by jointly deviating from the equilibrium. On the other hand, if it is locally uniformly stable, then the equilibrium must be weakly Pareto optimal. Moreover, we show that uniform stability determines the last-iterate convergence behavior for the family of incremental smoothed best-response dynamics, used to model individual and corporate behaviors in the markets. Unlike dynamics around strict equilibria, which can stabilize to socially-inefficient solutions, individually utility-seeking behaviors near mixed Nash equilibria lead to collective rationality.

Learnable Mixed Nash Equilibria are Collectively Rational

TL;DR

This work introduces uniform stability as a non-asymptotic stability notion for uncoupled learning dynamics in -player normal-form games and links it to a form of collective rationality via strategic Pareto optimality. It shows that local uniform stability implies strategic Pareto optimality and equivalence to pointwise uniform stability, while non-uniform uniform stability precludes convergence to mixed equilibria; these insights are tied to the spectral properties of the game Jacobian. The authors analyze incremental smoothed best-response dynamics, proving global convergence to -smoothed equilibria under locally uniformly stable regions and establishing a rate to Nash equilibria as , with non-convergence results when uniform stability fails. They extend the theory to partially mixed equilibria via reduced games and linear steepness assumptions, and discuss implications, limitations, and directions for future work, including broader non-asymptotic analyses and relaxation of interaction assumptions. Overall, the paper argues that while dynamics around strict equilibria can yield socially inefficient outcomes, dynamics near mixed Nash equilibria exhibit collective rationality governed by uniform stability, providing a refined lens on learnability and social outcomes in uncoupled settings.

Abstract

We extend the study of learning in games to dynamics that exhibit non-asymptotic stability. We do so through the notion of uniform stability, which is concerned with equilibria of individually utility-seeking dynamics. Perhaps surprisingly, it turns out to be closely connected to economic properties of collective rationality. Under mild non-degeneracy conditions and up to strategic equivalence, if a mixed equilibrium is not uniformly stable, then it is not weakly Pareto optimal: there is a way for all players to improve by jointly deviating from the equilibrium. On the other hand, if it is locally uniformly stable, then the equilibrium must be weakly Pareto optimal. Moreover, we show that uniform stability determines the last-iterate convergence behavior for the family of incremental smoothed best-response dynamics, used to model individual and corporate behaviors in the markets. Unlike dynamics around strict equilibria, which can stabilize to socially-inefficient solutions, individually utility-seeking behaviors near mixed Nash equilibria lead to collective rationality.
Paper Structure (42 sections, 37 theorems, 130 equations, 2 figures)

This paper contains 42 sections, 37 theorems, 130 equations, 2 figures.

Key Result

Lemma 3.1

Let $u, v \in \mathbb{R}^m \setminus \{0\}$. Then, $u^\top v > 0$ if and only if there is a positive-definite matrix $H$ so that:

Figures (2)

  • Figure 1: Uncoupled learning dynamics in two $2 \times 2$ normal-form games with purely-strategic utilities $\mathbf{f} = (f_1, f_2)$. The Nash equilibria are marked by stars. The streamlines visualize the trajectories of the learning dynamics. The heatmap plots the social welfare function $\min\{f_1, f_2\}$, measuring the utility of the player worst-off. The heatmap is white where the utilities are equal to the equilibrium; the darker the red, the worse the social welfare; the darker the blue, the better. (a) The mixed Nash equilibrium in this game is not weakly Pareto optimal, and the dynamics are unstable. (b) This equilibrium is weakly Pareto optimal, and the dynamics are non-asymptotically stable. The dynamics visualized here is continuous-time mirror ascent induced by the entropy mirror map.
  • Figure 2: The trajectories of $\beta$-smoothed best-response dynamics with $\eta$-learning rate toward a uniformly-stable, mixed Nash equilibrium (star) in a two-player normal-form game. (a) The $\beta$-smoothed equilibria become better approximations of the Nash equilibrium as $\beta$ shrinks, but the dynamics become less stable and exhibit more cycling. The figure on the left plots the trajectories initialized at the black dot for varying smoothing $\beta$ and fixed averaging $\eta$. (b) The learning dynamics converge to $\beta$-smoothed equilibria once $\eta$ becomes sufficiently small, and it does so with greater stability with smaller $\eta$. However, stability comes at the expense of slower rate of convergence. The figure on the right plots the trajectories for fixed $\beta$ and varying $\eta$.

Theorems & Definitions (93)

  • Definition 2.1: Normal-form game
  • Definition 2.2: Multilinear polynomial
  • Definition 2.3: Multilinear game
  • Definition 2.4: Nash equilibrium
  • Definition 2.5: Weak Pareto optimality
  • Definition 2.6: Strategic and non-strategic components
  • Definition 2.7: Strategic Pareto optimality
  • Definition 2.8: $\beta$-smoothed best-response
  • Definition 2.9: Steep regularizer
  • Definition 2.10: $\beta$-smoothed equilibrium
  • ...and 83 more