Table of Contents
Fetching ...

VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation

Son T. Luu, Trung Vo, Hiep Nguyen, Khanh Quoc Tran, Kiet Van Nguyen, Vu Tran, Ngan Luu-Thuy Nguyen, Le-Minh Nguyen

TL;DR

The VLSP 2025 MLQA-TSR paper introduces a Vietnamese multimodal legal question answering benchmark focused on traffic sign regulation, comprising retrieval (Subtask 1) and QA (Subtask 2). It provides a dataset construction methodology, including data collection from street traffic signs, annotation by multiple annotators, and a law database with two Vietnamese legal documents, plus evaluation protocols using $F2$-score for retrieval and accuracy for QA. Baseline methods and top-participant approaches leverage multimodal embeddings, vision-language LLMs, and prompting strategies (zero-shot, few-shot, chain-of-thought), with notable results: Subtask 1 best $F2$ = 64.55% and Subtask 2 best accuracy = 86.30%. The work demonstrates robust progress in Vietnamese multimodal legal processing, offers a public benchmark, and provides baseline code to spur further research in low-resource multilingual legal AI.

Abstract

This paper presents the VLSP 2025 MLQA-TSR - the multimodal legal question answering on traffic sign regulation shared task at VLSP 2025. VLSP 2025 MLQA-TSR comprises two subtasks: multimodal legal retrieval and multimodal question answering. The goal is to advance research on Vietnamese multimodal legal text processing and to provide a benchmark dataset for building and evaluating intelligent systems in multimodal legal domains, with a focus on traffic sign regulation in Vietnam. The best-reported results on VLSP 2025 MLQA-TSR are an F2 score of 64.55% for multimodal legal retrieval and an accuracy of 86.30% for multimodal question answering.

VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation

TL;DR

The VLSP 2025 MLQA-TSR paper introduces a Vietnamese multimodal legal question answering benchmark focused on traffic sign regulation, comprising retrieval (Subtask 1) and QA (Subtask 2). It provides a dataset construction methodology, including data collection from street traffic signs, annotation by multiple annotators, and a law database with two Vietnamese legal documents, plus evaluation protocols using -score for retrieval and accuracy for QA. Baseline methods and top-participant approaches leverage multimodal embeddings, vision-language LLMs, and prompting strategies (zero-shot, few-shot, chain-of-thought), with notable results: Subtask 1 best = 64.55% and Subtask 2 best accuracy = 86.30%. The work demonstrates robust progress in Vietnamese multimodal legal processing, offers a public benchmark, and provides baseline code to spur further research in low-resource multilingual legal AI.

Abstract

This paper presents the VLSP 2025 MLQA-TSR - the multimodal legal question answering on traffic sign regulation shared task at VLSP 2025. VLSP 2025 MLQA-TSR comprises two subtasks: multimodal legal retrieval and multimodal question answering. The goal is to advance research on Vietnamese multimodal legal text processing and to provide a benchmark dataset for building and evaluating intelligent systems in multimodal legal domains, with a focus on traffic sign regulation in Vietnam. The best-reported results on VLSP 2025 MLQA-TSR are an F2 score of 64.55% for multimodal legal retrieval and an accuracy of 86.30% for multimodal question answering.
Paper Structure (16 sections, 4 equations, 3 figures, 4 tables)

This paper contains 16 sections, 4 equations, 3 figures, 4 tables.

Figures (3)

  • Figure 1: A sample of a legal question about a traffic sign.
  • Figure 2: Data Creation Process.
  • Figure 3: Distribution of relevant articles in three sets