Learnable Mixed Nash Equilibria are Collectively Rational
Geelon So, Yi-An Ma
TL;DR
This work introduces uniform stability as a non-asymptotic stability notion for uncoupled learning dynamics in $N$-player normal-form games and links it to a form of collective rationality via strategic Pareto optimality. It shows that local uniform stability implies strategic Pareto optimality and equivalence to pointwise uniform stability, while non-uniform uniform stability precludes convergence to mixed equilibria; these insights are tied to the spectral properties of the game Jacobian. The authors analyze incremental smoothed best-response dynamics, proving global convergence to $eta$-smoothed equilibria under locally uniformly stable regions and establishing a $T^{-1/2}$ rate to Nash equilibria as $eta o 0$, with non-convergence results when uniform stability fails. They extend the theory to partially mixed equilibria via reduced games and linear steepness assumptions, and discuss implications, limitations, and directions for future work, including broader non-asymptotic analyses and relaxation of interaction assumptions. Overall, the paper argues that while dynamics around strict equilibria can yield socially inefficient outcomes, dynamics near mixed Nash equilibria exhibit collective rationality governed by uniform stability, providing a refined lens on learnability and social outcomes in uncoupled settings.
Abstract
We extend the study of learning in games to dynamics that exhibit non-asymptotic stability. We do so through the notion of uniform stability, which is concerned with equilibria of individually utility-seeking dynamics. Perhaps surprisingly, it turns out to be closely connected to economic properties of collective rationality. Under mild non-degeneracy conditions and up to strategic equivalence, if a mixed equilibrium is not uniformly stable, then it is not weakly Pareto optimal: there is a way for all players to improve by jointly deviating from the equilibrium. On the other hand, if it is locally uniformly stable, then the equilibrium must be weakly Pareto optimal. Moreover, we show that uniform stability determines the last-iterate convergence behavior for the family of incremental smoothed best-response dynamics, used to model individual and corporate behaviors in the markets. Unlike dynamics around strict equilibria, which can stabilize to socially-inefficient solutions, individually utility-seeking behaviors near mixed Nash equilibria lead to collective rationality.
