Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
Yu Wu, Ke Shu, Jonas Fischer, Lidia Pivovarova, David Rosson, Eetu Mäkelä, Mikko Tolonen
TL;DR
This work defines and benchmarks a novel multimodal task for detecting and extracting Latin fragments in 18th‑century historical books, using a ground-truth dataset of 724 ECCO pages with 12 Latin usage categories. It proposes a unified, model-agnostic prompt pipeline that operates on image, text, or both, and rigorously evaluates contemporary LLMs and a baseline linguistic detector under OCR noise through OCR post-correction and fuzzy token matching. Key findings show reliable Latin detection is achievable in zero-shot settings, with open-source LLMs reaching or surpassing baseline performance, particularly when multimodal input is leveraged; however, category-level analyses reveal limited semantic understanding and strong dependence on superficial textual cues. The study delivers a practical Latin detection pipeline, provides a valuable dataset for further research, and points to future work on semantic grounding, cross-domain generalization, and extension to additional historical languages.
Abstract
This paper presents a novel task of extracting Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal dataset of 724 annotated pages. The results demonstrate that reliable Latin detection with contemporary models is achievable. Our study provides the first comprehensive analysis of these models' capabilities and limits for this task.
