Autonomous Legged Mobile Manipulation for Lunar Surface Operations via Constrained Reinforcement Learning
Alvaro Belmonte-Baeza, Miguel Cazorla, Gabriel J. García, Carlos J. Pérez-Del-Pulgar, Jorge Pomares
TL;DR
This work addresses autonomous legged mobile manipulation for lunar surface operations by introducing a constrained reinforcement learning framework that jointly optimizes locomotion and manipulation while enforcing safety constraints under lunar gravity ($1/6\,g$). The method employs Constraints as Terminations (CaT) to integrate hard and soft safety limits into learning, enabling robust end-effector 6D pose tracking with a mean positional error of $4\,\text{cm}$ and orientation error of $8.1^{\circ}$, and strong constraint satisfaction. Experimental validation in NVIDIA Isaac Sim with domain randomization demonstrates both precise task execution and emergent energy-efficient behaviors in low gravity, highlighting the approach's potential for safe autonomous lunar exploration. The results bridge adaptive learning and mission-critical safety, advancing capabilities for integrated loco-manipulation on planetary surfaces and guiding future real-world validation and terrain challenges.
Abstract
Robotics plays a pivotal role in planetary science and exploration, where autonomous and reliable systems are crucial due to the risks and challenges inherent to space environments. The establishment of permanent lunar bases demands robotic platforms capable of navigating and manipulating in the harsh lunar terrain. While wheeled rovers have been the mainstay for planetary exploration, their limitations in unstructured and steep terrains motivate the adoption of legged robots, which offer superior mobility and adaptability. This paper introduces a constrained reinforcement learning framework designed for autonomous quadrupedal mobile manipulators operating in lunar environments. The proposed framework integrates whole-body locomotion and manipulation capabilities while explicitly addressing critical safety constraints, including collision avoidance, dynamic stability, and power efficiency, in order to ensure robust performance under lunar-specific conditions, such as reduced gravity and irregular terrain. Experimental results demonstrate the framework's effectiveness in achieving precise 6D task-space end-effector pose tracking, achieving an average positional accuracy of 4 cm and orientation accuracy of 8.1 degrees. The system consistently respects both soft and hard constraints, exhibiting adaptive behaviors optimized for lunar gravity conditions. This work effectively bridges adaptive learning with essential mission-critical safety requirements, paving the way for advanced autonomous robotic explorers for future lunar missions.
