Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring
Federico Di Gennaro, Khaled Eldowa, Nicolò Cesa-Bianchi
TL;DR
This work develops an instance-dependent regret analysis for nonstochastic linear partial monitoring with finite actions by introducing an Exploration-by-Optimization policy tailored to linear loss/observation structure. The approach yields minimax regret bounds that scale as $\widetilde{O}(\sqrt{T})$ in locally observable (easy) games and $\widetilde{O}(T^{2/3})$ in globally observable (hard) games, governed by interpretable alignment constants $\beta_{\text{loc}}$ and $\beta_{\text{glo}}$. The policy reduces to a convex optimization at each round, enabled by anchored loss estimators and an exponential weights update over Pareto-optimal actions, and it recovers tight performance in several classic settings such as full information and graph feedback. The paper also systematically instantiates the bounds across various partial-information scenarios, illustrating near-optimal behavior and highlighting practical computational tractability via semidefinite program representations for the per-round optimization.\n
Abstract
In contrast to the classic formulation of partial monitoring, linear partial monitoring can model infinite outcome spaces, while imposing a linear structure on both the losses and the observations. This setting can be viewed as a generalization of linear bandits where loss and feedback are decoupled in a flexible manner. In this work, we address a nonstochastic (adversarial), finite-actions version of the problem through a simple instance of the exploration-by-optimization method that is amenable to efficient implementation. We derive regret bounds that depend on the game structure in a more transparent manner than previous theoretical guarantees for this paradigm. Our bounds feature instance-specific quantities that reflect the degree of alignment between observations and losses, and resemble known guarantees in the stochastic setting. Notably, they achieve the standard $\sqrt{T}$ rate in easy (locally observable) games and $T^{2/3}$ in hard (globally observable) games, where $T$ is the time horizon. We instantiate these bounds in a selection of old and new partial information settings subsumed by this model, and illustrate that the achieved dependence on the game structure can be tight in interesting cases.
