Ensemble based Closed-Loop Optimal Control using Physics-Informed Neural Networks
Jostein Barry-Straume, Adwait D. Verulkar, Arash Sarshar, Andrey A. Popov, Adrian Sandu
TL;DR
This work develops an ensemble Physics-Informed Neural Network (PINN) framework to solve infinite-horizon, time-invariant optimal control problems by learning the optimal cost-to-go $\mathcal{J}$ and its costate $\lambda=\nabla_x \mathcal{J}$ from the Hamilton-Jacobi-Bellman equation. It introduces three ensemble-control strategies—individual, mean, and outlier-exclusion mean—showing that an ensemble with robust outlier handling can stabilize a two-state nonlinear affine system under noise and varied initial conditions without stabilizer terms. The methodology combines warm-start data with a dedicated HJB-PINN loss that enforces boundary conditions and PDE residuals, yielding accurate adjoints $\lambda$ and control signals $u^*$. The results demonstrate that the ensemble approach can reproduce analytical solutions and maintain stable closed-loop behavior, offering a scalable path toward more complex nonlinear control problems. The work contributes a practical, data-efficient framework for closed-loop control under uncertainty with potential applications in robotics and engineered systems where long-horizon optimality is essential.
Abstract
The objective of designing a control system is to steer a dynamical system with a control signal, guiding it to exhibit the desired behavior. The Hamilton-Jacobi-Bellman (HJB) partial differential equation offers a framework for optimal control system design. However, numerical solutions to this equation are computationally intensive, and analytical solutions are frequently unavailable. Knowledge-guided machine learning methodologies, such as physics-informed neural networks (PINNs), offer new alternative approaches that can alleviate the difficulties of solving the HJB equation numerically. This work presents a multistage ensemble framework to learn the optimal cost-to-go, and subsequently the corresponding optimal control signal, through the HJB equation. Prior PINN-based approaches rely on a stabilizing the HJB enforcement during training. Our framework does not use stabilizer terms and offers a means of controlling the nonlinear system, via either a singular learned control signal or an ensemble control signal policy. Success is demonstrated in closed-loop control, using both ensemble- and singular-control, of a steady-state time-invariant two-state continuous nonlinear system with an infinite time horizon, accounting of noisy, perturbed system states and varying initial conditions.
