Table of Contents
Fetching ...

Policy Gradient Method for LQG Control via Input-Output-History Representation: Convergence to $O(ε)$-Stationary Points

Tomonori Sadamoto, Takashi Tanaka

TL;DR

This paper tackles nonconvex policy optimization for LQG control using an input-output-history (IOH) representation. By showing an equivalence between dynamic output-feedback and static IOH gains on a finite history, and introducing a coerciveness-enhancing relaxation with covariance $\epsilon I$, it proves that a vanilla policy-gradient method converges to an $\mathcal{O}(\epsilon)$-stationary point of the original LQG cost. The approach yields a principled, gradient-based pathway to the LQG solution, with numerical experiments indicating recovery of the global optimum or near-optimal, including successful low-order controller synthesis via IOH. The work lays groundwork for model-free IOH-based LQG design and motivates future theoretical and algorithmic extensions toward global convergence and practical, data-driven implementations.

Abstract

We study the policy gradient method (PGM) for the linear quadratic Gaussian (LQG) dynamic output-feedback control problem using an input-output-history (IOH) representation of the closed-loop system. First, we show that any dynamic output-feedback controller is equivalent to a static partial-state feedback gain for a new system representation characterized by a finite-length IOH. Leveraging this equivalence, we reformulate the search for an optimal dynamic output feedback controller as an optimization problem over the corresponding partial-state feedback gain. Next, we introduce a relaxed version of the IOH-based LQG problem by incorporating a small process noise with covariance $εI$ into the new system to ensure coerciveness, a key condition for establishing gradient-based convergence guarantees. Consequently, we show that a vanilla PGM for the relaxed problem converges to an $\mathcal{O}(ε)$-stationary point, i.e., $\overline{K}$ satisfying $\|\nabla J(\overline{K})\|_F \leq \mathcal{O}(ε)$, where $J$ denotes the original LQG cost. Numerical experiments empirically indicate convergence to the vicinity of the globally optimal LQG controller.

Policy Gradient Method for LQG Control via Input-Output-History Representation: Convergence to $O(ε)$-Stationary Points

TL;DR

This paper tackles nonconvex policy optimization for LQG control using an input-output-history (IOH) representation. By showing an equivalence between dynamic output-feedback and static IOH gains on a finite history, and introducing a coerciveness-enhancing relaxation with covariance , it proves that a vanilla policy-gradient method converges to an -stationary point of the original LQG cost. The approach yields a principled, gradient-based pathway to the LQG solution, with numerical experiments indicating recovery of the global optimum or near-optimal, including successful low-order controller synthesis via IOH. The work lays groundwork for model-free IOH-based LQG design and motivates future theoretical and algorithmic extensions toward global convergence and practical, data-driven implementations.

Abstract

We study the policy gradient method (PGM) for the linear quadratic Gaussian (LQG) dynamic output-feedback control problem using an input-output-history (IOH) representation of the closed-loop system. First, we show that any dynamic output-feedback controller is equivalent to a static partial-state feedback gain for a new system representation characterized by a finite-length IOH. Leveraging this equivalence, we reformulate the search for an optimal dynamic output feedback controller as an optimization problem over the corresponding partial-state feedback gain. Next, we introduce a relaxed version of the IOH-based LQG problem by incorporating a small process noise with covariance into the new system to ensure coerciveness, a key condition for establishing gradient-based convergence guarantees. Consequently, we show that a vanilla PGM for the relaxed problem converges to an -stationary point, i.e., satisfying , where denotes the original LQG cost. Numerical experiments empirically indicate convergence to the vicinity of the globally optimal LQG controller.
Paper Structure (24 sections, 13 theorems, 105 equations, 4 figures, 1 algorithm)

This paper contains 24 sections, 13 theorems, 105 equations, 4 figures, 1 algorithm.

Key Result

Lemma 1

Suppose the system ${}^{\rm s}\space {\bm \Sigma}$ in hat Sigma is $L$-step observable, and let $z$ be the IOH of length $L$ defined by def_IOH, and $h$ be the full history defined by def_fullH. Consider a new system ${\bm \Sigma}$ defined by where $d$ is defined in def_d and If the initial state $h(L)$ of def_sigma coincides with the history of $\{u, y, w, v\}$ for $t \in \{0,\ldots, L-1\}$, th

Figures (4)

  • Figure 1: The colored solid lines show 20 variations of ${}^{\rm s}\space J({}^{\rm s}\space {\bm K}_i)$ for iteration $i$ when $L=3$ whereas the black dotted line shows ${}^{\rm s}\space J({}^{\rm s}\space {\bm K}_{\rm LQG})$.
  • Figure 2: The colored solid lines show the Bode diagrams of the 20 designed controllers ${}^{\rm s}\space {\bm K}_{10^5}$ when $L=3$. The black dotted line shows that of the true LQG controller ${}^{\rm s}\space {\bm K}_{\rm LQG}$.
  • Figure 3: The red solid line and blue dash-dotted line show the bode diagram of a designed four-dimensional ${}^{\rm s}\space {\bm K}_{10^5}$ and two-dimensional ${}^{\rm s}\space {\bm K}_{10^5}$, respectively, and the green dashed line shows that of ${}^{\rm s}\space {\bm K}_{\rm LQG}^{\rm red}$.
  • Figure 4: The colored lines show the four Hankel singular values of ${}^{\rm s}\space {\bm K}_{i}$ for iteration $i$ when $L=4$. Black dotted lines show the three Hankel singular values of ${}^{\rm s}\space {\bm K}_{\rm LQG}$.

Theorems & Definitions (31)

  • Definition 1
  • Definition 2
  • Lemma 1
  • proof
  • Lemma 2
  • proof
  • Lemma 3
  • proof
  • Theorem 1
  • proof
  • ...and 21 more