Trinity: A Modular Humanoid Robot AI System

Jingkai Sun; Qiang Zhang; Gang Han; Wen Zhao; Zhe Yong; Yan He; Jiaxu Wang; Jiahang Cao; Yijie Guo; Renjing Xu

Trinity: A Modular Humanoid Robot AI System

Jingkai Sun, Qiang Zhang, Gang Han, Wen Zhao, Zhe Yong, Yan He, Jiaxu Wang, Jiahang Cao, Yijie Guo, Renjing Xu

TL;DR

Trinity tackles the challenge of versatile humanoid robotics by unifying RL-based locomotion, visual-language perception, and language-driven task planning in a modular, hierarchical architecture. The locomotion module uses Adversarial Motion Priors with an AMP discriminator and a periodic/regularization reward structure within an MDP $(\mathcal{S},\mathcal{A},\mathcal{R},p,\gamma)$ and a policy $\pi(a_t|s_t)$ to achieve human-like motion; a Finite State Machine manages gait transitions. The perception module employs ManipVQA to fuse RGB-D visual data with semantic queries, producing actionable representations for the LLM planner, which in turn composes a sequence of robot skills from Arm, Hand, and Body capabilities with kinematic-aware prompting. Real-world experiments on a full-scale humanoid and safety-focused evaluations demonstrate robust loco-manipulation under dynamic upper-body movements, and safety constraints are enforced through the LLM-driven planning layer, enabling safer operation in unstructured environments. Overall, Trinity demonstrates the feasibility and benefits of an integrated, modular humanoid AI stack that leverages multimodal perception, long-horizon reasoning, and robust motion control to operate effectively in complex real-world settings.

Abstract

In recent years, research on humanoid robots has garnered increasing attention. With breakthroughs in various types of artificial intelligence algorithms, embodied intelligence, exemplified by humanoid robots, has been highly anticipated. The advancements in reinforcement learning (RL) algorithms have significantly improved the motion control and generalization capabilities of humanoid robots. Simultaneously, the groundbreaking progress in large language models (LLM) and visual language models (VLM) has brought more possibilities and imagination to humanoid robots. LLM enables humanoid robots to understand complex tasks from language instructions and perform long-term task planning, while VLM greatly enhances the robots' understanding and interaction with their environment. This paper introduces \textcolor{magenta}{Trinity}, a novel AI system for humanoid robots that integrates RL, LLM, and VLM. By combining these technologies, Trinity enables efficient control of humanoid robots in complex environments. This innovative approach not only enhances the capabilities but also opens new avenues for future research and applications of humanoid robotics.

Trinity: A Modular Humanoid Robot AI System

TL;DR

Abstract

Trinity: A Modular Humanoid Robot AI System

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (6)