Table of Contents
Fetching ...

Lyapunov-Aware Quantum-Inspired Reinforcement Learning for Continuous-Time Vehicle Control: A Feasibility Study

Nutkritta Kraipatthanapong, Natthaphat Thathong, Pannita Suksawas, Thanunnut Klunklin, Kritin Vongthonglua, Krit Attahakul, Aueaphum Aueawatthanaphisut

TL;DR

The paper tackles safe, provably stable control in continuous-time vehicle dynamics by integrating Lyapunov stability with a quantum-inspired reinforcement learning policy. It introduces LQRL, where a variational quantum circuit outputs acceleration commands and is constrained by a Lyapunov-based penalty to achieve asymptotic convergence. Through a longitudinal ACC case study, the approach demonstrates feasible safety integration and competitive performance relative to classical controllers, while revealing the need for adaptive Lyapunov gains to guarantee full stability in high-speed transients. The work provides a reproducible framework bridging quantum policy optimization and classical stability theory, with implications for quantum-safe control in autonomous systems and hybrid quantum–classical optimization.

Abstract

This paper presents a novel Lyapunov-Based Quantum Reinforcement Learning (LQRL) framework that integrates quantum policy optimization with Lyapunov stability analysis for continuous-time vehicle control. The proposed approach combines the representational power of variational quantum circuits (VQCs) with a stability-aware policy gradient mechanism to ensure asymptotic convergence and safe decision-making under dynamic environments. The vehicle longitudinal control problem was formulated as a continuous-state reinforcement learning task, where the quantum policy network generates control actions subject to Lyapunov stability constraints. Simulation experiments were conducted in a closed-loop adaptive cruise control scenario using a quantum-inspired policy trained under stability feedback. The results demonstrate that the LQRL framework successfully embeds Lyapunov stability verification into quantum policy learning, enabling interpretable and stability-aware control performance. Although transient overshoot and Lyapunov divergence were observed under aggressive acceleration, the system maintained bounded state evolution, validating the feasibility of integrating safety guarantees within quantum reinforcement learning architectures. The proposed framework provides a foundational step toward provably safe quantum control in autonomous systems and hybrid quantum-classical optimization domains.

Lyapunov-Aware Quantum-Inspired Reinforcement Learning for Continuous-Time Vehicle Control: A Feasibility Study

TL;DR

The paper tackles safe, provably stable control in continuous-time vehicle dynamics by integrating Lyapunov stability with a quantum-inspired reinforcement learning policy. It introduces LQRL, where a variational quantum circuit outputs acceleration commands and is constrained by a Lyapunov-based penalty to achieve asymptotic convergence. Through a longitudinal ACC case study, the approach demonstrates feasible safety integration and competitive performance relative to classical controllers, while revealing the need for adaptive Lyapunov gains to guarantee full stability in high-speed transients. The work provides a reproducible framework bridging quantum policy optimization and classical stability theory, with implications for quantum-safe control in autonomous systems and hybrid quantum–classical optimization.

Abstract

This paper presents a novel Lyapunov-Based Quantum Reinforcement Learning (LQRL) framework that integrates quantum policy optimization with Lyapunov stability analysis for continuous-time vehicle control. The proposed approach combines the representational power of variational quantum circuits (VQCs) with a stability-aware policy gradient mechanism to ensure asymptotic convergence and safe decision-making under dynamic environments. The vehicle longitudinal control problem was formulated as a continuous-state reinforcement learning task, where the quantum policy network generates control actions subject to Lyapunov stability constraints. Simulation experiments were conducted in a closed-loop adaptive cruise control scenario using a quantum-inspired policy trained under stability feedback. The results demonstrate that the LQRL framework successfully embeds Lyapunov stability verification into quantum policy learning, enabling interpretable and stability-aware control performance. Although transient overshoot and Lyapunov divergence were observed under aggressive acceleration, the system maintained bounded state evolution, validating the feasibility of integrating safety guarantees within quantum reinforcement learning architectures. The proposed framework provides a foundational step toward provably safe quantum control in autonomous systems and hybrid quantum-classical optimization domains.
Paper Structure (23 sections, 14 equations, 4 figures, 4 tables)

This paper contains 23 sections, 14 equations, 4 figures, 4 tables.

Figures (4)

  • Figure 1: Conceptual framework of the proposed Lyapunov-Based Quantum Reinforcement Learning (LQRL) applied to vehicle longitudinal dynamics. The QRL agent encodes state information into quantum circuits and updates the policy via gradient descent under Lyapunov stability constraints to ensure asymptotic safety.
  • Figure 2: System overview of the proposed LQRL framework integrating quantum policy, Lyapunov stability, and continuous vehicle dynamics.
  • Figure 3: Graphical simulation in Pygame showing the ego (blue) and lead (green) vehicles in the adaptive cruise control scenario.
  • Figure 4: Simulation results of the proposed Lyapunov-Based Quantum Reinforcement Learning (LQRL) framework in the adaptive cruise control task. (a) The spacing error $z(t)$ initially remained near zero and later exhibited a negative drift as the ego vehicle approached the lead vehicle, indicating a violation of the nominal headway constraint. (b) The velocity profiles show that the relative speed $v_r(t)$ diverged after 20 s, while the ego velocity $v_e(t)$ continued to increase under the upper acceleration bound $u_{\max}=3\,\text{m/s}^2$, leading to overshoot and transient instability. (c) The Lyapunov derivative $\dot{V}(x)$ was observed to become positive during this phase, revealing that the system energy increased rather than decayed, hence violating the Lyapunov decrease condition $\dot{V}(x)\le 0$. Collectively, these results demonstrate that while the LQRL agent learned a stability-aware policy structure, additional adaptive regularization or dynamic gain tuning is required to guarantee asymptotic stability during high-speed maneuvers.