Deep Reinforcement Learning for Sequential Combinatorial Auctions

Sai Srivatsa Ravindranath; Zhe Feng; Di Wang; Manzil Zaheer; Aranyak Mehta; David C. Parkes

Deep Reinforcement Learning for Sequential Combinatorial Auctions

Sai Srivatsa Ravindranath, Zhe Feng, Di Wang, Manzil Zaheer, Aranyak Mehta, David C. Parkes

TL;DR

This work tackles revenue optimization for Sequential Combinatorial Auctions (SCAs) where large, continuous action spaces hinder standard RL methods. It introduces a gradient-based fitted policy iteration that leverages differentiable transition dynamics and RochetNet-inspired menu structures with continuation-value offsets, enabling scalable learning up to 50 buyers and 50 items. The approach combines exact dynamic programming for small state spaces with neural actor-critic methods for larger spaces, and introduces entry-fee mechanisms to scale the method further. Empirical results show substantial revenue improvements over analytical baselines and PPO, with efficient training times and demonstrated scalability to large, realistic auction settings, bridging theory and practice in sequential auction design.

Abstract

Revenue-optimal auction design is a challenging problem with significant theoretical and practical implications. Sequential auction mechanisms, known for their simplicity and strong strategyproofness guarantees, are often limited by theoretical results that are largely existential, except for certain restrictive settings. Although traditional reinforcement learning methods such as Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) are applicable in this domain, they struggle with computational demands and convergence issues when dealing with large and continuous action spaces. In light of this and recognizing that we can model transitions differentiable for our settings, we propose using a new reinforcement learning framework tailored for sequential combinatorial auctions that leverages first-order gradients. Our extensive evaluations show that our approach achieves significant improvement in revenue over both analytical baselines and standard reinforcement learning algorithms. Furthermore, we scale our approach to scenarios involving up to 50 agents and 50 items, demonstrating its applicability in complex, real-world auction settings. As such, this work advances the computational tools available for auction design and contributes to bridging the gap between theoretical results and practical implementations in sequential auction design.

Deep Reinforcement Learning for Sequential Combinatorial Auctions

TL;DR

Abstract

Paper Structure (23 sections, 2 theorems, 6 equations, 4 figures, 5 tables, 4 algorithms)

This paper contains 23 sections, 2 theorems, 6 equations, 4 figures, 5 tables, 4 algorithms.

Introduction
Main Challenges.
Our Contributions.
Related Works.
Preliminaries
Method
Exact Method for Small Number of States
Approximate Methods for Large Number of States
Neural Network Architecture.
Fitted Policy Iteration.
Entry Fee Mechanisms for Extremely Large Number of States
Experimental Results
Constrained Additive Valuations
Combinatorial Valuations
Scaling up
...and 8 more sections

Key Result

Proposition 0

For a current policy $\pi$ and value function $V_\pi(.)$, the improved policy $\pi'$ for a state $s^t$ is given by:

Figures (4)

Figure 1: Fitted Policy Iteration (FPI)
Figure : Sequential Combinatorial Mechanisms with Menus
Figure : Policy Improvement Step
Figure : Dynamic Program (DP)

Theorems & Definitions (5)

Remark 0
Proposition 0
Remark 0
Proposition 0
proof

Deep Reinforcement Learning for Sequential Combinatorial Auctions

TL;DR

Abstract

Deep Reinforcement Learning for Sequential Combinatorial Auctions

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (4)

Theorems & Definitions (5)