VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation
Son T. Luu, Trung Vo, Hiep Nguyen, Khanh Quoc Tran, Kiet Van Nguyen, Vu Tran, Ngan Luu-Thuy Nguyen, Le-Minh Nguyen
TL;DR
The VLSP 2025 MLQA-TSR paper introduces a Vietnamese multimodal legal question answering benchmark focused on traffic sign regulation, comprising retrieval (Subtask 1) and QA (Subtask 2). It provides a dataset construction methodology, including data collection from street traffic signs, annotation by multiple annotators, and a law database with two Vietnamese legal documents, plus evaluation protocols using $F2$-score for retrieval and accuracy for QA. Baseline methods and top-participant approaches leverage multimodal embeddings, vision-language LLMs, and prompting strategies (zero-shot, few-shot, chain-of-thought), with notable results: Subtask 1 best $F2$ = 64.55% and Subtask 2 best accuracy = 86.30%. The work demonstrates robust progress in Vietnamese multimodal legal processing, offers a public benchmark, and provides baseline code to spur further research in low-resource multilingual legal AI.
Abstract
This paper presents the VLSP 2025 MLQA-TSR - the multimodal legal question answering on traffic sign regulation shared task at VLSP 2025. VLSP 2025 MLQA-TSR comprises two subtasks: multimodal legal retrieval and multimodal question answering. The goal is to advance research on Vietnamese multimodal legal text processing and to provide a benchmark dataset for building and evaluating intelligent systems in multimodal legal domains, with a focus on traffic sign regulation in Vietnam. The best-reported results on VLSP 2025 MLQA-TSR are an F2 score of 64.55% for multimodal legal retrieval and an accuracy of 86.30% for multimodal question answering.
