Delay as Payoff in MAB

Ofir Schlisselberg; Ido Cohen; Tal Lancewicki; Yishay Mansour

Delay as Payoff in MAB

Ofir Schlisselberg, Ido Cohen, Tal Lancewicki, Yishay Mansour

TL;DR

This paper investigates a variant of the classical stochastic Multi-armed Bandit problem, where the payoff received by an agent is both delayed, and directly corresponds to the magnitude of the delay.

Abstract

In this paper, we investigate a variant of the classical stochastic Multi-armed Bandit (MAB) problem, where the payoff received by an agent (either cost or reward) is both delayed, and directly corresponds to the magnitude of the delay. This setting models faithfully many real world scenarios such as the time it takes for a data packet to traverse a network given a choice of route (where delay serves as the agent's cost); or a user's time spent on a web page given a choice of content (where delay serves as the agent's reward). Our main contributions are tight upper and lower bounds for both the cost and reward settings. For the case that delays serve as costs, which we are the first to consider, we prove optimal regret that scales as $\sum_{i:Δ_i > 0}\frac{\log T}{Δ_i} + d^*$, where $T$ is the maximal number of steps, $Δ_i$ are the sub-optimality gaps and $d^*$ is the minimal expected delay amongst arms. For the case that delays serves as rewards, we show optimal regret of $\sum_{i:Δ_i > 0}\frac{\log T}{Δ_i} + \bar{d}$, where $\bar d$ is the second maximal expected delay. These improve over the regret in the general delay-dependent payoff setting, which scales as $\sum_{i:Δ_i > 0}\frac{\log T}{Δ_i} + D$, where $D$ is the maximum possible delay. Our regret bounds highlight the difference between the cost and reward scenarios, showing that the improvement in the cost scenario is more significant than for the reward. Finally, we accompany our theoretical results with an empirical evaluation.

Delay as Payoff in MAB

TL;DR

Abstract

, where

is the maximal number of steps,

are the sub-optimality gaps and

is the minimal expected delay amongst arms. For the case that delays serves as rewards, we show optimal regret of

, where

is the second maximal expected delay. These improve over the regret in the general delay-dependent payoff setting, which scales as

, where

is the maximum possible delay. Our regret bounds highlight the difference between the cost and reward scenarios, showing that the improvement in the cost scenario is more significant than for the reward. Finally, we accompany our theoretical results with an empirical evaluation.

Paper Structure (32 sections, 37 theorems, 109 equations, 1 figure, 1 table, 10 algorithms)

This paper contains 32 sections, 37 theorems, 109 equations, 1 figure, 1 table, 10 algorithms.

Introduction
Our Contributions
Cost vs Reward - intuition
Paper organization
Related work
Problem Setup
Delay as cost
CSE Algorithm
case (i):
case (ii):
case (iii):
Bounded Doubling Successive Elimination
Lower Bound
Conservative SE algorithms:
Delay as Reward
...and 17 more sections

Key Result

Lemma 4.1

For every step $t$, if the last $\min {\left\{ D, t \right\}}$ steps was played with a round robin of a set of size at least $R$:

Figures (1)

Figure 1: This graph shows results of experiments on different algorithms (color) and different distributions (line style).

Theorems & Definitions (41)

Lemma 4.1
Definition 4.2
Lemma 4.3
Lemma 4.4: Safe Elimination
Theorem 4.5
Corollary 4.6
Theorem 4.7
Theorem 5.1: Safe Elimination
Theorem 5.2
Corollary 5.3
...and 31 more

Delay as Payoff in MAB

TL;DR

Abstract

Delay as Payoff in MAB

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (1)

Theorems & Definitions (41)