RoleMotion: A Large-Scale Dataset towards Robust Scene-Specific Role-Playing Motion Synthesis with Fine-grained Descriptions
Junran Peng, Yiheng Huang, Silei Shen, Zeji Wei, Jingwei Yang, Baojie Wang, Yonghao He, Chuanchen Luo, Man Zhang, Xucheng Yin, Wei Sui
TL;DR
RoleMotion tackles a core bottleneck in text-to-motion research by offering a large-scale, scene-specific motion dataset captured from scratch with fine-grained textual descriptions. It pairs the data with a strong transformer-based evaluator and a new benchmark to align text prompts with body-and-hand motions, enabling rigorous comparison across methods. The study demonstrates that two-stage body-and-hand generation can improve textual alignment and motion quality, and provides qualitative and quantitative evidence that RoleMotion outperforms existing datasets in realism and controllability. This work delivers practical resources and evaluation tools to advance robust scene-specific motion synthesis for NPCs in virtual environments.
Abstract
In this paper, we introduce RoleMotion, a large-scale human motion dataset that encompasses a wealth of role-playing and functional motion data tailored to fit various specific scenes. Existing text datasets are mainly constructed decentrally as amalgamation of assorted subsets that their data are nonfunctional and isolated to work together to cover social activities in various scenes. Also, the quality of motion data is inconsistent, and textual annotation lacks fine-grained details in these datasets. In contrast, RoleMotion is meticulously designed and collected with a particular focus on scenes and roles. The dataset features 25 classic scenes, 110 functional roles, over 500 behaviors, and 10296 high-quality human motion sequences of body and hands, annotated with 27831 fine-grained text descriptions. We build an evaluator stronger than existing counterparts, prove its reliability, and evaluate various text-to-motion methods on our dataset. Finally, we explore the interplay of motion generation of body and hands. Experimental results demonstrate the high-quality and functionality of our dataset on text-driven whole-body generation.
