CoSER: Coordinating LLM-Based Persona Simulation of Established Roles

Xintao Wang; Heng Wang; Yifei Zhang; Xinfeng Yuan; Rui Xu; Jen-tse Huang; Siyu Yuan; Haoran Guo; Jiangjie Chen; Shuchang Zhou; Wei Wang; Yanghua Xiao

CoSER: Coordinating LLM-Based Persona Simulation of Established Roles

Xintao Wang, Heng Wang, Yifei Zhang, Xinfeng Yuan, Rui Xu, Jen-tse Huang, Siyu Yuan, Haoran Guo, Jiangjie Chen, Shuchang Zhou, Wei Wang, Yanghua Xiao

TL;DR

CoSER tackles the core challenges of authentic data scarcity and evaluation bias in role-playing LLMs for established characters by delivering a large, multi-type dataset derived from 771 books and introducing the given-circumstance acting framework for training and evaluation. It trains open models CoSER 8B and 70B on a broad, diverse corpus and demonstrates state-of-the-art performance on multiple benchmarks, including human evaluation, through a penalty-based, rubric-driven GCA protocol. The work also demonstrates the value of retrieval augmentation and inner thoughts in improving fidelity and control, and provides extensive analyses, ablation studies, and case studies to validate the approach. The authors plan to release the dataset, models, and evaluation tools to support further research while addressing copyright and ethical considerations.

Abstract

Role-playing language agents (RPLAs) have emerged as promising applications of large language models (LLMs). However, simulating established characters presents a challenging task for RPLAs, due to the lack of authentic character datasets and nuanced evaluation methods using such data. In this paper, we present CoSER, a collection of a high-quality dataset, open models, and an evaluation protocol towards effective RPLAs of established characters. The CoSER dataset covers 17,966 characters from 771 renowned books. It provides authentic dialogues with real-world intricacies, as well as diverse data types such as conversation setups, character experiences and internal thoughts. Drawing from acting methodology, we introduce given-circumstance acting for training and evaluating role-playing LLMs, where LLMs sequentially portray multiple characters in book scenes. Using our dataset, we develop CoSER 8B and CoSER 70B, i.e., advanced open role-playing LLMs built on LLaMA-3.1 models. Extensive experiments demonstrate the value of the CoSER dataset for RPLA training, evaluation and retrieval. Moreover, CoSER 70B exhibits state-of-the-art performance surpassing or matching GPT-4o on our evaluation and three existing benchmarks, i.e., achieving 75.80% and 93.47% accuracy on the InCharacter and LifeChoice benchmarks respectively.

CoSER: Coordinating LLM-Based Persona Simulation of Established Roles

TL;DR

Abstract

CoSER: Coordinating LLM-Based Persona Simulation of Established Roles

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (6)