On the Optimal Regret of Locally Private Linear Contextual Bandit

Jiachun Li; David Simchi-Levi; Yining Wang

On the Optimal Regret of Locally Private Linear Contextual Bandit

Jiachun Li, David Simchi-Levi, Yining Wang

TL;DR

This work addresses the challenge of designing regret-optimal, locally private linear contextual bandits. It proves that $\tilde{O}(\sqrt{T})$ regret is achievable under adaptive local privacy constraints by moving beyond mean-squared-error analyses and adopting mean-absolute deviation (MAD) based guarantees. The authors introduce Layered Private Linear Regression (LPLR), an adaptive input-partitioning framework that partitions the context-action input into hierarchical bins and performs per-bin PCA regression with privacy-preserving noise. By embedding LPLR as a locally private offline regression oracle within an action-elimination RegCB framework, they derive near-optimal regret bounds that scale as $\tilde{O}(\sqrt{T})$ up to dimension-dependent factors, and discuss the exponential dependence on dimension as a key open challenge. The results have implications for privacy-aware bandits and potentially broader privacy-constrained learning tasks, suggesting new directions for reducing dimension dependence and extending to broader model classes such as GLMs and large action spaces.

Abstract

Contextual bandit with linear reward functions is among one of the most extensively studied models in bandit and online learning research. Recently, there has been increasing interest in designing \emph{locally private} linear contextual bandit algorithms, where sensitive information contained in contexts and rewards is protected against leakage to the general public. While the classical linear contextual bandit algorithm admits cumulative regret upper bounds of $\tilde O(\sqrt{T})$ via multiple alternative methods, it has remained open whether such regret bounds are attainable in the presence of local privacy constraints, with the state-of-the-art result being $\tilde O(T^{3/4})$. In this paper, we show that it is indeed possible to achieve an $\tilde O(\sqrt{T})$ regret upper bound for locally private linear contextual bandit. Our solution relies on several new algorithmic and analytical ideas, such as the analysis of mean absolute deviation errors and layered principal component regression in order to achieve small mean absolute deviation errors.

On the Optimal Regret of Locally Private Linear Contextual Bandit

TL;DR

This work addresses the challenge of designing regret-optimal, locally private linear contextual bandits. It proves that

regret is achievable under adaptive local privacy constraints by moving beyond mean-squared-error analyses and adopting mean-absolute deviation (MAD) based guarantees. The authors introduce Layered Private Linear Regression (LPLR), an adaptive input-partitioning framework that partitions the context-action input into hierarchical bins and performs per-bin PCA regression with privacy-preserving noise. By embedding LPLR as a locally private offline regression oracle within an action-elimination RegCB framework, they derive near-optimal regret bounds that scale as

up to dimension-dependent factors, and discuss the exponential dependence on dimension as a key open challenge. The results have implications for privacy-aware bandits and potentially broader privacy-constrained learning tasks, suggesting new directions for reducing dimension dependence and extending to broader model classes such as GLMs and large action spaces.

Abstract

via multiple alternative methods, it has remained open whether such regret bounds are attainable in the presence of local privacy constraints, with the state-of-the-art result being

. In this paper, we show that it is indeed possible to achieve an

regret upper bound for locally private linear contextual bandit. Our solution relies on several new algorithmic and analytical ideas, such as the analysis of mean absolute deviation errors and layered principal component regression in order to achieve small mean absolute deviation errors.

Paper Structure (36 sections, 15 theorems, 136 equations, 2 figures, 6 algorithms)

This paper contains 36 sections, 15 theorems, 136 equations, 2 figures, 6 algorithms.

Introduction
Model specification and definitions
Related works
Motivating negative results and examples
Fundamental limits of the mean-square error
Fundamental limits of input perturbation
Discussion: necessity of input partitioning
Action elimination framework
Locally private offline regression oracle with confidence intervals
The action elimination algorithm
Layered private linear regression (LPLR)
Key idea: Adaptive layered partitioning of $d$-dimensional space
Overview of the algorithm
The LPLR-Update subroutine.
The LPLR-PCR subroutine.
...and 21 more sections

Key Result

Theorem 1

Fix $\alpha\in(0,1]$ and $d=2$. There exists a numerical constant $\underline C_1>0$ such that the following holds: for any $n\geq 1$, there exists a distribution $\mathcal{D}_n$ supported on $\{x\in\mathbb R^d: \|x\|_2\leq 1\}$ and parameter class $\Theta_n\subseteq\{\theta\in\mathbb R^d:\|\theta\| where $\Lambda_{\mathcal{D}_n}=\mathbb E_{\mathcal{D}_n}[\phi\phi^\top]$.

Figures (2)

Figure 1: Illustration of an example layered partitioning of $3$-dimensional space
Figure 2: Iterative procedure of partitioning $(\phi,y)$

Theorems & Definitions (25)

Definition 1
Theorem 1
Theorem 2
Definition 2: Locally private offline oracle with confidence intervals
Theorem 3
Proposition 1: Local privacy of LPLR-Update
Lemma 1: Utility of LPLR-Update
Lemma 2
Lemma 3
Remark 1
...and 15 more

On the Optimal Regret of Locally Private Linear Contextual Bandit

TL;DR

Abstract

On the Optimal Regret of Locally Private Linear Contextual Bandit

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (2)

Theorems & Definitions (25)