VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage
A. Alfarano, L. Venturoli, D. Negueruela del Castillo
TL;DR
VQArt-Bench tackles the gap in deep semantic evaluation for art-focused VQA by introducing an agent-based question-generation pipeline that creates semantically rich, linguistically diverse questions grounded in art descriptions. The benchmark spans seven visual reasoning dimensions and comprises 14,463 multiple-choice items, evaluated across 14 state-of-the-art MLLMs to reveal how current systems handle symbolic meaning, narratives, and complex visual relationships in cultural heritage. Key findings show widespread limitations, including a surprising weakness in counting tasks and a consistent advantage for closed-source models over open-source ones, while larger models generally perform better. The work demonstrates the value of semantically grounded, domain-specific benchmarks for advancing true visual literacy in cultural heritage VQA and provides a scalable, controllable pipeline for future benchmark enrichment.
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in joint visual and linguistic tasks. However, existing Visual Question Answering (VQA) benchmarks often fail to evaluate deep semantic understanding, particularly in complex domains like visual art analysis. Confined to simple syntactic structures and surface-level attributes, these questions fail to capture the diversity and depth of human visual inquiry. This limitation incentivizes models to exploit statistical shortcuts rather than engage in visual reasoning. To address this gap, we introduce VQArt-Bench, a new, large-scale VQA benchmark for the cultural heritage domain. This benchmark is constructed using a novel multi-agent pipeline where specialized agents collaborate to generate nuanced, validated, and linguistically diverse questions. The resulting benchmark is structured along relevant visual understanding dimensions that probe a model's ability to interpret symbolic meaning, narratives, and complex visual relationships. Our evaluation of 14 state-of-the-art MLLMs on this benchmark reveals significant limitations in current models, including a surprising weakness in simple counting tasks and a clear performance gap between proprietary and open-source models.
