Actor-Free Continuous Control via Structurally Maximizable Q-Functions
Yigit Korkmaz, Urvi Bhuwania, Ayush Jain, Erdem Bıyık
TL;DR
This work addresses the challenge of performing value-based reinforcement learning in continuous action spaces without an actor. It introduces Q3C, which uses a structurally maximizable Q-function built from control-points, with an action-conditioned Q-value generator and a wire-fitting interpolator to identify the maximizing action directly. Through relevance-based filtering, control-point diversification, and scale-aware normalization, Q3C achieves competitive or superior performance to state-of-the-art actor-critic baselines, especially in environments with restricted action spaces. The results, supported by ablations and visualizations, demonstrate robustness and stability across diverse tasks, and the authors release code to facilitate adoption and further research.
Abstract
Value-based algorithms are a cornerstone of off-policy reinforcement learning due to their simplicity and training stability. However, their use has traditionally been restricted to discrete action spaces, as they rely on estimating Q-values for individual state-action pairs. In continuous action spaces, evaluating the Q-value over the entire action space becomes computationally infeasible. To address this, actor-critic methods are typically employed, where a critic is trained on off-policy data to estimate Q-values, and an actor is trained to maximize the critic's output. Despite their popularity, these methods often suffer from instability during training. In this work, we propose a purely value-based framework for continuous control that revisits structural maximization of Q-functions, introducing a set of key architectural and algorithmic choices to enable efficient and stable learning. We evaluate the proposed actor-free Q-learning approach on a range of standard simulation tasks, demonstrating performance and sample efficiency on par with state-of-the-art baselines, without the cost of learning a separate actor. Particularly, in environments with constrained action spaces, where the value functions are typically non-smooth, our method with structural maximization outperforms traditional actor-critic methods with gradient-based maximization. We have released our code at https://github.com/USC-Lira/Q3C.
