Sample-Based Hybrid Mode Control: Asymptotically Optimal Switching of Algorithmic and Non-Differentiable Control Modes
Yilang Liu, Haoxiang You, Ian Abraham
TL;DR
The paper tackles the problem of dynamically switching between discrete, potentially non-differentiable and algorithmic control modes in robotics by formulating an indefinite hybrid mode switching problem and introducing a discrete-time, sample-based optimization approach. It redefines the problem through a discrete-time horizon with decisions over the mode, start time, and duration $(m,\tau,\lambda)$, and then develops an iterative, sample-based method that draws uniform candidates from the single-mode transition set $\Omega$ to achieve asymptotic convergence guarantees. Key contributions include a discrete-time definite formulation, an iterative sequencing scheme with local optimality properties, and a scalable, GPU-friendly, sample-based solver that can handle long-horizon tasks and high-dimensional state spaces, demonstrated on a Cartpole and a high-dimensional quadruped with real hardware. The approach enables complex mode compositions by integrating algorithmic controllers with learned policies (e.g., MPC/MPPI and learned feedback controllers) and has practical impact for enabling agile, stable, and reactive behaviors in legged robots and similar systems.
Abstract
This paper investigates a sample-based solution to the hybrid mode control problem across non-differentiable and algorithmic hybrid modes. Our approach reasons about a set of hybrid control modes as an integer-based optimization problem where we select what mode to apply, when to switch to another mode, and the duration for which we are in a given control mode. A sample-based variation is derived to efficiently search the integer domain for optimal solutions. We find our formulation yields strong performance guarantees that can be applied to a number of robotics-related tasks. In addition, our approach is able to synthesize complex algorithms and policies to compound behaviors and achieve challenging tasks. Last, we demonstrate the effectiveness of our approach in real-world robotic examples that require reactive switching between long-term planning and high-frequency control.
