Table of Contents
Fetching ...

VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage

A. Alfarano, L. Venturoli, D. Negueruela del Castillo

TL;DR

VQArt-Bench tackles the gap in deep semantic evaluation for art-focused VQA by introducing an agent-based question-generation pipeline that creates semantically rich, linguistically diverse questions grounded in art descriptions. The benchmark spans seven visual reasoning dimensions and comprises 14,463 multiple-choice items, evaluated across 14 state-of-the-art MLLMs to reveal how current systems handle symbolic meaning, narratives, and complex visual relationships in cultural heritage. Key findings show widespread limitations, including a surprising weakness in counting tasks and a consistent advantage for closed-source models over open-source ones, while larger models generally perform better. The work demonstrates the value of semantically grounded, domain-specific benchmarks for advancing true visual literacy in cultural heritage VQA and provides a scalable, controllable pipeline for future benchmark enrichment.

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in joint visual and linguistic tasks. However, existing Visual Question Answering (VQA) benchmarks often fail to evaluate deep semantic understanding, particularly in complex domains like visual art analysis. Confined to simple syntactic structures and surface-level attributes, these questions fail to capture the diversity and depth of human visual inquiry. This limitation incentivizes models to exploit statistical shortcuts rather than engage in visual reasoning. To address this gap, we introduce VQArt-Bench, a new, large-scale VQA benchmark for the cultural heritage domain. This benchmark is constructed using a novel multi-agent pipeline where specialized agents collaborate to generate nuanced, validated, and linguistically diverse questions. The resulting benchmark is structured along relevant visual understanding dimensions that probe a model's ability to interpret symbolic meaning, narratives, and complex visual relationships. Our evaluation of 14 state-of-the-art MLLMs on this benchmark reveals significant limitations in current models, including a surprising weakness in simple counting tasks and a clear performance gap between proprietary and open-source models.

VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage

TL;DR

VQArt-Bench tackles the gap in deep semantic evaluation for art-focused VQA by introducing an agent-based question-generation pipeline that creates semantically rich, linguistically diverse questions grounded in art descriptions. The benchmark spans seven visual reasoning dimensions and comprises 14,463 multiple-choice items, evaluated across 14 state-of-the-art MLLMs to reveal how current systems handle symbolic meaning, narratives, and complex visual relationships in cultural heritage. Key findings show widespread limitations, including a surprising weakness in counting tasks and a consistent advantage for closed-source models over open-source ones, while larger models generally perform better. The work demonstrates the value of semantically grounded, domain-specific benchmarks for advancing true visual literacy in cultural heritage VQA and provides a scalable, controllable pipeline for future benchmark enrichment.

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in joint visual and linguistic tasks. However, existing Visual Question Answering (VQA) benchmarks often fail to evaluate deep semantic understanding, particularly in complex domains like visual art analysis. Confined to simple syntactic structures and surface-level attributes, these questions fail to capture the diversity and depth of human visual inquiry. This limitation incentivizes models to exploit statistical shortcuts rather than engage in visual reasoning. To address this gap, we introduce VQArt-Bench, a new, large-scale VQA benchmark for the cultural heritage domain. This benchmark is constructed using a novel multi-agent pipeline where specialized agents collaborate to generate nuanced, validated, and linguistically diverse questions. The resulting benchmark is structured along relevant visual understanding dimensions that probe a model's ability to interpret symbolic meaning, narratives, and complex visual relationships. Our evaluation of 14 state-of-the-art MLLMs on this benchmark reveals significant limitations in current models, including a surprising weakness in simple counting tasks and a clear performance gap between proprietary and open-source models.
Paper Structure (26 sections, 4 figures, 2 tables)

This paper contains 26 sections, 4 figures, 2 tables.

Figures (4)

  • Figure 1: Examples from our VQArt-bench with highlighted related subjects for better visualization. VQArt-bench is deeply grounded in the related images and refers to specific elements or areas of the artwork. Each question requires a profound visual understanding to be answered correctly.
  • Figure 2: Demonstration of question quality improvement using our agentic pipeline (Ours) over a rule-based system (AQUA). All the questions have been generated from the same data source. While the rule-based questions are often shallow, lack nuance, and can be factually inconsistent, our method produces context-aware questions appropriate for fine-grained analysis. The rule-based approach produces semantically shallow questions that lack correct terminology (e.g., referring to the Madonna as a "woman") and can introduce factual inaccuracies, such as hallucinating an "animal on the shirt". In contrast, our agentic pipeline leverages LLMs to generate questions that are both challenging and precise. It preserves the nuances of the source material, formulating sophisticated questions that require a deep understanding of artistic composition (left example) and complex symbolism (right example). Correct answers are reported in bold text.
  • Figure 3: Examples from our VQArt-Bench dataset, categorized by our seven core evaluation dimensions. Our benchmark is designed to test a spectrum of visual reasoning skills. It challenges models on foundational abilities like identifying objects and their properties (Instance Identity, Instance Attribute), locating them in the scene (Instance Location), and quantifying them (Instance Counting). The evaluation then progresses to more complex compositional tasks, such as understanding Spatial Relation and Instance Interaction, and high-level tasks that require inferring context and causality (Visual-Inspired Reasoning). Correct answers are highlighted in Bold text.
  • Figure 4: Exploitative LLM based statistical evaluation of key attributes in our VQArt-Bench. The dataset shows broad diversity in terms of compositional elements, including (a) a wide range of human figure counts, from individuals to large crowds; (b) varied numbers of symbolic objects per scene; and (f) representation across different genders. The collection spans multiple genres and settings, covering (c) primary artistic genres like religious and portraiture; (d, e) a balance of scenes with and without animal or vegetation presence; and (g, h) a mix of indoor and diverse outdoor environments. Finally, the dataset captures rich atmospheric and stylistic variations, including (l) different times of day; (i) various weather conditions; (k) simple and complex lighting sources; and (j) a full spectrum of warm-to-cool color ratios, indicating diverse visual moods.