Competitive Multi-armed Bandit Games for Resource Sharing

Hongbo Li; Lingjie Duan

Competitive Multi-armed Bandit Games for Resource Sharing

Hongbo Li, Lingjie Duan

TL;DR

This work studies N-player competitive MAB (CMAB) games for resource sharing under unknown Bernoulli rewards with collisions. It develops threshold-based policies for both selfish and socially optimal play, reveals that selfish behavior can cause an infinite price of anarchy ($PoA$), and shows that informational mechanisms alone cannot fix this when players are non-myopic. It then introduces the Combined Informational and Side-Payment (CISP) mechanism, which yields $PoA=1$ by aligning incentives through information sharing and monetary transfers while preserving budget balance. Experiments corroborate the theory, demonstrating that CISP matches the convergence pace of the social optimum and eliminates the inefficiency observed under selfish behavior and information hiding.

Abstract

In modern resource-sharing systems, multiple agents access limited resources with unknown stochastic conditions to perform tasks. When multiple agents access the same resource (arm) simultaneously, they compete for successful usage, leading to contention and reduced rewards. This motivates our study of competitive multi-armed bandit (CMAB) games. In this paper, we study a new N-player K-arm competitive MAB game, where non-myopic players (agents) compete with each other to form diverse private estimations of unknown arms over time. Their possible collisions on same arms and time-varying nature of arm rewards make the policy analysis more involved than existing studies for myopic players. We explicitly analyze the threshold-based structures of social optimum and existing selfish policy, showing that the latter causes prolonged convergence time $Ω(\frac{K}{η^2}\ln({\frac{KN}δ}))$, while socially optimal policy with coordinated communication reduces it to $\mathcal{O}(\frac{K}{Nη^2}\ln{(\frac{K}δ)})$. Based on the comparison, we prove that the competition among selfish players for the best arm can result in an infinite price of anarchy (PoA), indicating an arbitrarily large efficiency loss compared to social optimum. We further prove that no informational (non-monetary) mechanism (including Bayesian persuasion) can reduce the infinite PoA, as the strategic misreporting by non-myopic players undermines such approaches. To address this, we propose a Combined Informational and Side-Payment (CISP) mechanism, which provides socially optimal arm recommendations with proper informational and monetary incentives to players according to their time-varying private beliefs. Our CISP mechanism keeps ex-post budget balanced for social planner and ensures truthful reporting from players, achieving the minimum PoA=1 and same convergence time as social optimum.

Competitive Multi-armed Bandit Games for Resource Sharing

TL;DR

), and shows that informational mechanisms alone cannot fix this when players are non-myopic. It then introduces the Combined Informational and Side-Payment (CISP) mechanism, which yields

by aligning incentives through information sharing and monetary transfers while preserving budget balance. Experiments corroborate the theory, demonstrating that CISP matches the convergence pace of the social optimum and eliminates the inefficiency observed under selfish behavior and information hiding.

Abstract

, while socially optimal policy with coordinated communication reduces it to

. Based on the comparison, we prove that the competition among selfish players for the best arm can result in an infinite price of anarchy (PoA), indicating an arbitrarily large efficiency loss compared to social optimum. We further prove that no informational (non-monetary) mechanism (including Bayesian persuasion) can reduce the infinite PoA, as the strategic misreporting by non-myopic players undermines such approaches. To address this, we propose a Combined Informational and Side-Payment (CISP) mechanism, which provides socially optimal arm recommendations with proper informational and monetary incentives to players according to their time-varying private beliefs. Our CISP mechanism keeps ex-post budget balanced for social planner and ensures truthful reporting from players, achieving the minimum PoA=1 and same convergence time as social optimum.

Competitive Multi-armed Bandit Games for Resource Sharing

TL;DR

Abstract

Competitive Multi-armed Bandit Games for Resource Sharing

TL;DR

Abstract

Paper Structure

Table of Contents

Key Result

Figures (4)

Theorems & Definitions (17)