Table of Contents
Fetching ...

Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics

Sunmook Choi, Yahya Sattar, Yassir Jedra, Maryam Fazel, Sarah Dean

TL;DR

A nonstationary bandit problem where rewards depend on both actions and latent states, the latter governed by unknown linear dynamics is studied, providing a sub-optimality guarantee for this problem, enabling the regret upper bound.

Abstract

We study a nonstationary bandit problem where rewards depend on both actions and latent states, the latter governed by unknown linear dynamics. Crucially, the state dynamics also depend on the actions, resulting in tension between short-term and long-term rewards. We propose an explore-then-commit algorithm for a finite horizon $T$. During the exploration phase, random Rademacher actions enable estimation of the Markov parameters of the linear dynamics, which characterize the action-reward relationship. In the commit phase, the algorithm uses the estimated parameters to design an optimized action sequence for long-term reward. Our proposed algorithm achieves $\tilde{\mathcal{O}}(T^{2/3})$ regret. Our analysis handles two key challenges: learning from temporally correlated rewards, and designing action sequences with optimal long-term reward. We address the first challenge by providing near-optimal sample complexity and error bounds for system identification using bilinear rewards. We address the second challenge by proving an equivalence with indefinite quadratic optimization over a hypercube, a known NP-hard problem. We provide a sub-optimality guarantee for this problem, enabling our regret upper bound. Lastly, we propose a semidefinite relaxation with Goemans-Williamson rounding as a practical approach.

Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics

TL;DR

A nonstationary bandit problem where rewards depend on both actions and latent states, the latter governed by unknown linear dynamics is studied, providing a sub-optimality guarantee for this problem, enabling the regret upper bound.

Abstract

We study a nonstationary bandit problem where rewards depend on both actions and latent states, the latter governed by unknown linear dynamics. Crucially, the state dynamics also depend on the actions, resulting in tension between short-term and long-term rewards. We propose an explore-then-commit algorithm for a finite horizon . During the exploration phase, random Rademacher actions enable estimation of the Markov parameters of the linear dynamics, which characterize the action-reward relationship. In the commit phase, the algorithm uses the estimated parameters to design an optimized action sequence for long-term reward. Our proposed algorithm achieves regret. Our analysis handles two key challenges: learning from temporally correlated rewards, and designing action sequences with optimal long-term reward. We address the first challenge by providing near-optimal sample complexity and error bounds for system identification using bilinear rewards. We address the second challenge by proving an equivalence with indefinite quadratic optimization over a hypercube, a known NP-hard problem. We provide a sub-optimality guarantee for this problem, enabling our regret upper bound. Lastly, we propose a semidefinite relaxation with Goemans-Williamson rounding as a practical approach.
Paper Structure (44 sections, 20 theorems, 153 equations, 5 figures, 1 algorithm)

This paper contains 44 sections, 20 theorems, 153 equations, 5 figures, 1 algorithm.

Key Result

Proposition 1

Let ${\boldsymbol{S}}_T {:=} {\boldsymbol{M}}_T {+} {\boldsymbol{M}}_T^\top$. Then, the optimal open-loop action sequence ${\boldsymbol{u}}_{0:T}^\star$ is the solution to the problem:

Figures (5)

  • Figure 1: Graphical Model for Non-Stationary Bandits with controlled Latent Dynamics.
  • Figure 2: Each curve shows a mean over 20 experiments, with shaded regions indicating $\pm1$ standard deviation. (a) Expected cumulative reward under the oracle benchmark, approximated by semidefinite relaxation with Goemans-Williamson rounding (SDP+GW) and by the sign-iteration method (SignIter). (b) The regret of the explore-then-commit algorithm measured against the SDP+GW oracle benchmark, compared with the theoretical $\tilde{\mathcal{O}}(T^{2/3})$ rate. (c)--(d) Relative error of Markov parameter estimation for different truncation lengths $L$, under systems with spectral radii $\rho({\boldsymbol{A}})=0.1$ and $\rho({\boldsymbol{A}})=0.9$.
  • Figure 3: Comparison between SDP+GW and SignIter to the true optimum from the brute force method.
  • Figure 4: Expected cumulative reward under the oracle benchmark, approximated by semidefinite relaxation and Goemans-Williamson rounding (SDP+GW) and by the sign-iteration method (SignIter), in different spectral radii $\rho({\boldsymbol{A}})$.
  • Figure 5: The regret curves of the explore-then-commit algorithm measured from SDP+GW and SignIter against the SDP+GW oracle benchmark, compared with the theoretical $\tilde{\mathcal{O}}(T^{2/3})$ rate.

Theorems & Definitions (36)

  • Proposition 1
  • Proposition 2
  • Theorem 3
  • Theorem 4
  • Proposition 5
  • Proposition 6
  • Proposition 2
  • proof
  • Definition 1: Persistence of Excitation kumar2015stochastic
  • Theorem 3: Persistence of Excitation sattar2025learning
  • ...and 26 more