MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
Sathyanarayanan Ramamoorthy, Vishwa Shah, Simran Khanuja, Zaid Sheikh, Shan Jie, Ann Chia, Shearman Chua, Graham Neubig
TL;DR
MERLIN presents a novel multilingual multimodal entity linking benchmark built from BBC news titles paired with images across five languages. It formalizes MMEL, curates a 1,000-title-per-language dataset annotated via Prolific and INCEpTION, and provides baselines with mGENRE and GEMEL leveraging Llama-2 and Aya-23 encoders. Experiments demonstrate that visual context consistently improves linking accuracy, especially for ambiguous mentions and languages with limited multilingual training, with English translations offering additional gains. The 공개 데이터와 방법들은 멀티링구얼 멀티모달 EL 연구를 촉진하는 벤치마크로 활용될 수 있다.
Abstract
This paper introduces MERLIN, a novel testbed system for the task of Multilingual Multimodal Entity Linking. The created dataset includes BBC news article titles, paired with corresponding images, in five languages: Hindi, Japanese, Indonesian, Vietnamese, and Tamil, featuring over 7,000 named entity mentions linked to 2,500 unique Wikidata entities. We also include several benchmarks using multilingual and multimodal entity linking methods exploring different language models like LLaMa-2 and Aya-23. Our findings indicate that incorporating visual data improves the accuracy of entity linking, especially for entities where the textual context is ambiguous or insufficient, and particularly for models that do not have strong multilingual abilities. For the work, the dataset, methods are available here at https://github.com/rsathya4802/merlin
