Table of Contents
Fetching ...

No-Regret Online Autobidding Algorithms in First-price Auctions

Yuan Deng, Yilin Li, Wei Tang, Hanrui Zhang

TL;DR

The paper addresses no-regret autobidding in repeated first-price auctions with ROI constraints, focusing on nontruthful mechanisms and two feedback models. It develops a mirror-descent–style framework that reduces randomized optimal strategies to deterministic bidding via a concave envelope and a convexified rival bid distribution, enabling staged bootstrap with distribution learning. The main results show near-optimal regret bounds: $ ilde{O}( oot 4 extersh{T})$ in full feedback and $ ilde{O}(T^{3/4})$ in bandit feedback, along with ROI violation guarantees, thereby advancing theory and practice for ROI-constrained autobidders in nontruthful settings. These contributions offer practical, provably robust bidding strategies for advertisers and platform designers facing nontruthful auction mechanisms and uncertainty about rival bids.

Abstract

Automated bidding to optimize online advertising with various constraints, e.g. ROI constraints and budget constraints, is widely adopted by advertisers. A key challenge lies in designing algorithms for non-truthful mechanisms with ROI constraints. While prior work has addressed truthful auctions or non-truthful auctions with weaker benchmarks, this paper provides a significant improvement: We develop online bidding algorithms for repeated first-price auctions with ROI constraints, benchmarking against the optimal randomized strategy in hindsight. In the full feedback setting, where the maximum competing bid is observed, our algorithm achieves a near-optimal $\widetilde{O}(\sqrt{T})$ regret bound, and in the bandit feedback setting (where the bidder only observes whether the bidder wins each auction), our algorithm attains $\widetilde{O}(T^{3/4})$ regret bound.

No-Regret Online Autobidding Algorithms in First-price Auctions

TL;DR

The paper addresses no-regret autobidding in repeated first-price auctions with ROI constraints, focusing on nontruthful mechanisms and two feedback models. It develops a mirror-descent–style framework that reduces randomized optimal strategies to deterministic bidding via a concave envelope and a convexified rival bid distribution, enabling staged bootstrap with distribution learning. The main results show near-optimal regret bounds: in full feedback and in bandit feedback, along with ROI violation guarantees, thereby advancing theory and practice for ROI-constrained autobidders in nontruthful settings. These contributions offer practical, provably robust bidding strategies for advertisers and platform designers facing nontruthful auction mechanisms and uncertainty about rival bids.

Abstract

Automated bidding to optimize online advertising with various constraints, e.g. ROI constraints and budget constraints, is widely adopted by advertisers. A key challenge lies in designing algorithms for non-truthful mechanisms with ROI constraints. While prior work has addressed truthful auctions or non-truthful auctions with weaker benchmarks, this paper provides a significant improvement: We develop online bidding algorithms for repeated first-price auctions with ROI constraints, benchmarking against the optimal randomized strategy in hindsight. In the full feedback setting, where the maximum competing bid is observed, our algorithm achieves a near-optimal regret bound, and in the bandit feedback setting (where the bidder only observes whether the bidder wins each auction), our algorithm attains regret bound.
Paper Structure (33 sections, 20 theorems, 161 equations, 3 algorithms)

This paper contains 33 sections, 20 theorems, 161 equations, 3 algorithms.

Key Result

Theorem 3.1

Fixing any input value sequence $V^T$. For any competing bid distribution $F$, there exists a distribution $F_{\mathrm{conv}}\in\Delta([0, 1])$ such that: For any (randomized) strategy under $V^T$ and $F$, there is a deterministic strategy under $V^T$ and $F_{\mathrm{conv}}$ inducing reward and paym

Theorems & Definitions (43)

  • Definition 3.1: Allocation-payment curve $G$
  • Example 3.2: Suboptimality of deterministic strategy
  • Theorem 3.1: Optimal randomized bidding strategy
  • proof : Proof of \ref{['thm:optimal_randomized_strategy']}
  • Proposition 3.2
  • Theorem 4.1
  • proof : Proof of \ref{['thm:randomized_algorithm_full_feedback']}
  • Theorem 5.1
  • Lemma A.1
  • proof : Proof of \ref{['lem:monoton allocPay']}
  • ...and 33 more