Table of Contents
Fetching ...

Data-driven learning of feedback maps for explicit robust predictive control: an approximation theoretic view

Siddhartha Ganguly, Shubham Gupta, Debasish Chatterjee

TL;DR

The paper addresses robust MPC for uncertain linear systems by marrying data-driven learning with explicit policy synthesis. It solves the underlying minmax OCP exactly at grid points via a convex SIP solved with the MSA algorithm to generate state–action data, then learns explicit feedback maps with uniform error guarantees using QuIFS for low dimensions and NNFS for moderate dimensions. The authors prove that the learned policies preserve recursive feasibility and exhibit ISS-like stability, and they demonstrate two numerical examples showing improved region of attraction and fast online evaluation relative to traditional multiparametric methods. The approach offers a practical, provably reliable path to real-time robust control with data-driven explicit controllers, applicable to systems where online optimization is too costly.

Abstract

We establish an algorithm to learn feedback maps from data for a class of robust model predictive control (MPC) problems. The algorithm accounts for the approximation errors due to the learning directly at the synthesis stage, ensuring recursive feasibility by construction. The optimal control problem consists of a linear noisy dynamical system, a quadratic stage and quadratic terminal costs as the objective, and convex constraints on the state, control, and disturbance sequences; the control minimizes and the disturbance maximizes the objective. We proceed via two steps -- (a) Data generation: First, we reformulate the given minmax problem into a convex semi-infinite program and employ recently developed tools to solve it in an exact fashion on grid points of the state space to generate (state, action) data. (b) Learning approximate feedback maps: We employ a couple of approximation schemes that furnish tight approximations within preassigned uniform error bounds on the admissible state space to learn the unknown feedback policy. The stability of the closed-loop system under the approximate feedback policies is also guaranteed under a standard set of hypotheses. Two benchmark numerical examples are provided to illustrate the results.

Data-driven learning of feedback maps for explicit robust predictive control: an approximation theoretic view

TL;DR

The paper addresses robust MPC for uncertain linear systems by marrying data-driven learning with explicit policy synthesis. It solves the underlying minmax OCP exactly at grid points via a convex SIP solved with the MSA algorithm to generate state–action data, then learns explicit feedback maps with uniform error guarantees using QuIFS for low dimensions and NNFS for moderate dimensions. The authors prove that the learned policies preserve recursive feasibility and exhibit ISS-like stability, and they demonstrate two numerical examples showing improved region of attraction and fast online evaluation relative to traditional multiparametric methods. The approach offers a practical, provably reliable path to real-time robust control with data-driven explicit controllers, applicable to systems where online optimization is too costly.

Abstract

We establish an algorithm to learn feedback maps from data for a class of robust model predictive control (MPC) problems. The algorithm accounts for the approximation errors due to the learning directly at the synthesis stage, ensuring recursive feasibility by construction. The optimal control problem consists of a linear noisy dynamical system, a quadratic stage and quadratic terminal costs as the objective, and convex constraints on the state, control, and disturbance sequences; the control minimizes and the disturbance maximizes the objective. We proceed via two steps -- (a) Data generation: First, we reformulate the given minmax problem into a convex semi-infinite program and employ recently developed tools to solve it in an exact fashion on grid points of the state space to generate (state, action) data. (b) Learning approximate feedback maps: We employ a couple of approximation schemes that furnish tight approximations within preassigned uniform error bounds on the admissible state space to learn the unknown feedback policy. The stability of the closed-loop system under the approximate feedback policies is also guaranteed under a standard set of hypotheses. Two benchmark numerical examples are provided to illustrate the results.
Paper Structure (18 sections, 7 theorems, 57 equations, 7 figures, 4 tables)

This paper contains 18 sections, 7 theorems, 57 equations, 7 figures, 4 tables.

Key Result

lemma 1

Consider the OCP eq:approx ready param robust MPC and the SIP eq:approx ready param SIP robust MPC with their associated data and notations. Define the set of admissible $(\theta, \eta, r)$ by Then $\feasSip$ is closed and convex.

Figures (7)

  • Figure 2: The superimposed regions of attraction --- computed using YALMIP (violet) and Algorithm \ref{['alg:msap']} (magenta) --- are shown in the left-hand subfigure. The right-hand subfigure displays sample phase-space trajectories starting from the initial state $\xz \Let (1\,\,1)^{\top}$, obtained using both Algorithm \ref{['alg:msap']} and YALMIP’s robust optimization module.
  • Figure 3: The left-hand and the right-hand subfigures depicts the learned policy $\mutrunc(\cdot)$ computed via Algorithm \ref{['alg:extension_algo']} and the policy $\mu^*_0(\cdot)$, respectively.
  • Figure 4: The left-hand subfigure shows that the optimal value consistently converges to the same value across multiple runs of simulated annealing when Algorithm \ref{['alg:msap']} is employed. The right-hand subfigure depicts the error surface corresponding to the feedback policies in Figure \ref{['fig:uopt_uapp_rmpc_ex_1']}; it can be seen that the prespecified error margin $\eps=0.03$ is respected
  • Figure 5: The approximate policy $\mutrunc(\cdot)$ (the left-hand subfigure) obtained via Algorithm \ref{['alg: nn_algo']} and the corresponding error surface (the right-hand subfigure).
  • Figure 6: Validation loss of the NN during training.
  • ...and 2 more figures

Theorems & Definitions (14)

  • remark 1
  • lemma 1
  • theorem 1
  • remark 2
  • remark 3
  • theorem 2
  • remark 4
  • theorem 3
  • remark 5
  • definition 1
  • ...and 4 more