The Confusing Instance Principle for Online Linear Quadratic Control
Waris Radji, Odalric-Ambrym Maillard
TL;DR
This paper introduces the Confusing Instance (CI) principle as a new lens for exploration in online Linear Quadratic Regulation with unknown dynamics. It extends the Minimum Empirical Divergence (MED) framework from discrete settings to continuous LQR via MED-LQ, combining rank-one perturbations, stability constraints, and an empirical-divergence objective to drive exploration toward informative model perturbations. MED-LQ avoids confidence-bound based methods, leverages explicit LQR structure and Lyapunov stability, and demonstrates competitive performance on classical control benchmarks and industrial-like tasks, including auto-stabilization scenarios. The work lays a theoretical and algorithmic foundation for CI-based exploration in large-scale MDPs and suggests future directions for formal regret analysis and deep RL extensions with CI-guided exploration.
Abstract
We revisit the problem of controlling linear systems with quadratic cost under unknown dynamics with model-based reinforcement learning. Traditional methods like Optimism in the Face of Uncertainty and Thompson Sampling, rooted in multi-armed bandits (MABs), face practical limitations. In contrast, we propose an alternative based on the Confusing Instance (CI) principle, which underpins regret lower bounds in MABs and discrete Markov Decision Processes (MDPs) and is central to the Minimum Empirical Divergence (MED) family of algorithms, known for their asymptotic optimality in various settings. By leveraging the structure of LQR policies along with sensitivity and stability analysis, we develop MED-LQ. This novel control strategy extends the principles of CI and MED beyond small-scale settings. Our benchmarks on a comprehensive control suite demonstrate that MED-LQ achieves competitive performance in various scenarios while highlighting its potential for broader applications in large-scale MDPs.
