RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models

Yeongtak Oh; Dohyun Chung; Juhyeon Shin; Sangha Park; Johan Barthelemy; Jisoo Mok; Sungroh Yoon

RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models

Yeongtak Oh, Dohyun Chung, Juhyeon Shin, Sangha Park, Johan Barthelemy, Jisoo Mok, Sungroh Yoon

TL;DR

RePIC introduces a reinforcement-learning-based post-training framework to personalize multimodal language models for image captioning. It leverages verifiable rewards—Object Consistency Tuning (OCT), Visual Localization Tuning (VLT), and Identity Consistency Tuning (ICT)—within a Group Relative Policy Optimization (GRPO) scheme to enhance both visual recognition and personalized generation, reducing reliance on large-scale high-quality captions. Empirical results show substantial gains over SFT-based baselines, especially in challenging multi-concept scenarios, while maintaining general captioning capabilities. The approach represents a data-efficient path to robust real-world personalization for MLLMs with potential impact on personalized assistants and accessible AI systems.

Abstract

Recent multi-modal large language models (MLLMs) often struggle to generate personalized image captions, even when trained on high-quality captions. In this work, we observe that such limitations persist in existing post-training-based MLLM personalization methods. Specifically, despite being post-tuned with large-scale caption data through supervised fine-tuning (SFT), these models frequently fail to produce faithful descriptions in real-world scenarios, such as multi-concept image captioning. However, acquiring large-scale, high-quality captions for such complex settings is both costly and difficult. To address the data-centric nature of SFT, we propose a reinforcement learning (RL)-based post-training framework. To the best of our knowledge, this is the first RL-based approach to post-train MLLMs for personalized image captioning. Our method significantly enhances both visual recognition and personalized generation capabilities of MLLMs, and consistently outperforms existing SFT-based baselines, especially in the challenging multi-concept image captioning task. Project page: https://github.com/oyt9306/RePIC

RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models

TL;DR

Abstract

RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (23)