Learning to Design Soft Hands using Reward Models
Xueqian Bai, Nicklas Hansen, Adabhav Singh, Michael T. Tolley, Yan Duan, Pieter Abbeel, Xiaolong Wang, Sha Yi
TL;DR
Soft hands enable safe interaction but designing them to be both compliant and functional is challenging due to high-dimensional morphology–control coupling and costly evaluations. The paper presents CEM-RM, a data-driven framework that learns a design distribution over the variables $l^{fle}$, $l^{seg}$, $h_i$, $\mathbf{p}$, and $\psi$, and accelerates search by coupling Cross-Entropy Method with a neural reward model trained from teleoperation-grounded simulation rewards. Contributions include a parallel FEM-based soft-hand design space, a data-efficient optimization loop that reduces environment interactions by more than half, and hardware validation showing improved grasp success across diverse objects; ablation insights on teleoperation data and population size. Significance lies in providing a practical blueprint for data-driven co-design in soft robotics, enabling robust, rapid discovery of tendon-driven soft-hand morphologies and guiding future work on sim-to-real transfer and integrated control-design optimization.
Abstract
Soft robotic hands promise to provide compliant and safe interaction with objects and environments. However, designing soft hands to be both compliant and functional across diverse use cases remains challenging. Although co-design of hardware and control better couples morphology to behavior, the resulting search space is high-dimensional, and even simulation-based evaluation is computationally expensive. In this paper, we propose a Cross-Entropy Method with Reward Model (CEM-RM) framework that efficiently optimizes tendon-driven soft robotic hands based on teleoperation control policy, reducing design evaluations by more than half compared to pure optimization while learning a distribution of optimized hand designs from pre-collected teleoperation data. We derive a design space for a soft robotic hand composed of flexural soft fingers and implement parallelized training in simulation. The optimized hands are then 3D-printed and deployed in the real world using both teleoperation data and real-time teleoperation. Experiments in both simulation and hardware demonstrate that our optimized design significantly outperforms baseline hands in grasping success rates across a diverse set of challenging objects.
