Table of Contents
Fetching ...

Document Intelligence in the Era of Large Language Models: A Survey

Weishi Wang, Hengchang Hu, Zhijie Zhang, Zhaochen Li, Hongxin Shao, Daniel Dahlmeier

TL;DR

The paper surveys the emergence of document AI powered by decoder-only LLMs, outlining how multimodal and multilingual capabilities are reshaping tasks such as KIE, DLA, DC, DS, QA, and DCG. It compares prompt-based and unified-encoding strategies, reviews retrieval-augmented approaches, and discusses document-specific foundation models and agent-based DocAgent futures. Key findings include state-of-the-art methods like LayoutLLM, InstructDoc, mPLUG-DocOwl, UDOP, RAPTOR, and VisRAG, and a common need for better cross-lingual generalization, layout understanding, and long-context efficiency. The work highlights practical implications for scalable, robust, and multilingual DAI systems and guides future research on collaborative agents and domain-aware foundation models.

Abstract

Document AI (DAI) has emerged as a vital application area, and is significantly transformed by the advent of large language models (LLMs). While earlier approaches relied on encoder-decoder architectures, decoder-only LLMs have revolutionized DAI, bringing remarkable advancements in understanding and generation. This survey provides a comprehensive overview of DAI's evolution, highlighting current research attempts and future prospects of LLMs in this field. We explore key advancements and challenges in multimodal, multilingual, and retrieval-augmented DAI, while also suggesting future research directions, including agent-based approaches and document-specific foundation models. This paper aims to provide a structured analysis of the state-of-the-art in DAI and its implications for both academic and practical applications.

Document Intelligence in the Era of Large Language Models: A Survey

TL;DR

The paper surveys the emergence of document AI powered by decoder-only LLMs, outlining how multimodal and multilingual capabilities are reshaping tasks such as KIE, DLA, DC, DS, QA, and DCG. It compares prompt-based and unified-encoding strategies, reviews retrieval-augmented approaches, and discusses document-specific foundation models and agent-based DocAgent futures. Key findings include state-of-the-art methods like LayoutLLM, InstructDoc, mPLUG-DocOwl, UDOP, RAPTOR, and VisRAG, and a common need for better cross-lingual generalization, layout understanding, and long-context efficiency. The work highlights practical implications for scalable, robust, and multilingual DAI systems and guides future research on collaborative agents and domain-aware foundation models.

Abstract

Document AI (DAI) has emerged as a vital application area, and is significantly transformed by the advent of large language models (LLMs). While earlier approaches relied on encoder-decoder architectures, decoder-only LLMs have revolutionized DAI, bringing remarkable advancements in understanding and generation. This survey provides a comprehensive overview of DAI's evolution, highlighting current research attempts and future prospects of LLMs in this field. We explore key advancements and challenges in multimodal, multilingual, and retrieval-augmented DAI, while also suggesting future research directions, including agent-based approaches and document-specific foundation models. This paper aims to provide a structured analysis of the state-of-the-art in DAI and its implications for both academic and practical applications.
Paper Structure (41 sections, 1 table)