Table of Contents
Fetching ...

Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring

Federico Di Gennaro, Khaled Eldowa, Nicolò Cesa-Bianchi

TL;DR

This work develops an instance-dependent regret analysis for nonstochastic linear partial monitoring with finite actions by introducing an Exploration-by-Optimization policy tailored to linear loss/observation structure. The approach yields minimax regret bounds that scale as $\widetilde{O}(\sqrt{T})$ in locally observable (easy) games and $\widetilde{O}(T^{2/3})$ in globally observable (hard) games, governed by interpretable alignment constants $\beta_{\text{loc}}$ and $\beta_{\text{glo}}$. The policy reduces to a convex optimization at each round, enabled by anchored loss estimators and an exponential weights update over Pareto-optimal actions, and it recovers tight performance in several classic settings such as full information and graph feedback. The paper also systematically instantiates the bounds across various partial-information scenarios, illustrating near-optimal behavior and highlighting practical computational tractability via semidefinite program representations for the per-round optimization.\n

Abstract

In contrast to the classic formulation of partial monitoring, linear partial monitoring can model infinite outcome spaces, while imposing a linear structure on both the losses and the observations. This setting can be viewed as a generalization of linear bandits where loss and feedback are decoupled in a flexible manner. In this work, we address a nonstochastic (adversarial), finite-actions version of the problem through a simple instance of the exploration-by-optimization method that is amenable to efficient implementation. We derive regret bounds that depend on the game structure in a more transparent manner than previous theoretical guarantees for this paradigm. Our bounds feature instance-specific quantities that reflect the degree of alignment between observations and losses, and resemble known guarantees in the stochastic setting. Notably, they achieve the standard $\sqrt{T}$ rate in easy (locally observable) games and $T^{2/3}$ in hard (globally observable) games, where $T$ is the time horizon. We instantiate these bounds in a selection of old and new partial information settings subsumed by this model, and illustrate that the achieved dependence on the game structure can be tight in interesting cases.

Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring

TL;DR

This work develops an instance-dependent regret analysis for nonstochastic linear partial monitoring with finite actions by introducing an Exploration-by-Optimization policy tailored to linear loss/observation structure. The approach yields minimax regret bounds that scale as in locally observable (easy) games and in globally observable (hard) games, governed by interpretable alignment constants and . The policy reduces to a convex optimization at each round, enabled by anchored loss estimators and an exponential weights update over Pareto-optimal actions, and it recovers tight performance in several classic settings such as full information and graph feedback. The paper also systematically instantiates the bounds across various partial-information scenarios, illustrating near-optimal behavior and highlighting practical computational tractability via semidefinite program representations for the per-round optimization.\n

Abstract

In contrast to the classic formulation of partial monitoring, linear partial monitoring can model infinite outcome spaces, while imposing a linear structure on both the losses and the observations. This setting can be viewed as a generalization of linear bandits where loss and feedback are decoupled in a flexible manner. In this work, we address a nonstochastic (adversarial), finite-actions version of the problem through a simple instance of the exploration-by-optimization method that is amenable to efficient implementation. We derive regret bounds that depend on the game structure in a more transparent manner than previous theoretical guarantees for this paradigm. Our bounds feature instance-specific quantities that reflect the degree of alignment between observations and losses, and resemble known guarantees in the stochastic setting. Notably, they achieve the standard rate in easy (locally observable) games and in hard (globally observable) games, where is the time horizon. We instantiate these bounds in a selection of old and new partial information settings subsumed by this model, and illustrate that the achieved dependence on the game structure can be tight in interesting cases.
Paper Structure (36 sections, 35 theorems, 224 equations, 1 figure, 3 algorithms)

This paper contains 36 sections, 35 theorems, 224 equations, 1 figure, 3 algorithms.

Key Result

Proposition 0

Assume that the game has at least two non-duplicate Pareto optimal actions, and that $\mathcal{B}_2(r) \subseteq \mathcal{L}$ for some $r>0$. Then, the minimax regret $R^*_T$ is $\widetilde{\Omega}(\sqrt{T})$ if the game is locally observable, $\widetilde{\Omega}(T^{2/3})$ if it is globally but not

Figures (1)

  • Figure 1: The first figure represents a (non-zero) assignment of weights on a $9$-node cycle such that the sum at any three consecutive nodes is $0$, illustrating that $\bm{M}$ is singular. On a 10-node cycle, the second figure provides a representation of $\bm{M}^{-1} e_a$; i.e., the $a$-th column of $\bm{M}^{-1}$, where $a$ is an arbitrary action in $\mathcal{A}$. The last figure is a representation of $\bm{M}^{-1} (e_a-e_b)$ on the same graph, where $b$ is a neighbor of $a$.

Theorems & Definitions (61)

  • Definition 1: Observability conditions kirschner2020information
  • Proposition 0
  • Lemma 0
  • Lemma 0
  • Proposition 1
  • Lemma 0
  • Lemma 0
  • Theorem 0
  • Theorem 0
  • Lemma 1
  • ...and 51 more