EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture
Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad, Hesham Abdelgawad, Mohamed Aref, Ahmed Fares
TL;DR
EgMM-Corpus addresses the scarcity of culturally grounded multimodal data for Egyptian culture by introducing a dedicated dataset with 313 concepts and ~3,130 images across landmarks, food, and folklore, each with textual grounding from Wikipedia/Britannica and visual sources. The authors design a modular data collection pipeline that automatically retrieves concepts from TasteAtlas and UNESCO, gathers textual backgrounds, and builds per-concept directories with images and background descriptions, all under ethical and quality controls. They provide a baseline evaluation using zero-shot CLIP, reporting Top-1 21.2% and Top-5 36.4% classification accuracy and cross-modal retrieval challenges (I2T R@1 21.2%, I2T R@5 27.2%), highlighting existing cultural biases in large vision-language models. The work establishes EgMM-Corpus as a practical benchmark for developing culturally aware vision-language models and outlines future expansion to more regions, inclusion of VQA tasks via LLMs, and multimodal reasoning benchmarks.
Abstract
Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularly for regions in the Middle East and Africa. In this paper, we introduce EgMM-Corpus, a multimodal dataset dedicated to Egyptian culture. By designing and running a new data collection pipeline, we collected over 3,000 images, covering 313 concepts across landmarks, food, and folklore. Each entry in the dataset is manually validated for cultural authenticity and multimodal coherence. EgMM-Corpus aims to provide a reliable resource for evaluating and training vision-language models in an Egyptian cultural context. We further evaluate the zero-shot performance of Contrastive Language-Image Pre-training CLIP on EgMM-Corpus, on which it achieves 21.2% Top-1 accuracy and 36.4% Top-5 accuracy in classification. These results underscore the existing cultural bias in large-scale vision-language models and demonstrate the importance of EgMM-Corpus as a benchmark for developing culturally aware models.
