Table of Contents
Fetching ...

LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding

Sheikh Jubair, Arwa Omayrah, Amal Alshammari, Alhanoof Althnian, Abdulhamed Alothaimen, Norah A. Alzahrani, Shahad D. Alzaidi, Nora Al-Twairesh, Abdulmohsen Al-Thubaity

TL;DR

LC-Eval introduces a bilingual multi-task benchmark to assess long-context understanding in English and Arabic, covering 4K to over 128K tokens across four tasks: multi-document QA, bilingual QA, claim verification, and MCQ QA. The authors augment the dataset with an entity-relationship–based evaluation that uses LLMs as semantic judges, and validate all data with multiple human annotators. The dataset comprises 7,903 samples drawn from diverse public sources and includes careful curation to stress deep reasoning, information tracing, and cross-lingual extraction. Experimental results show that even strong models like GPT-4o struggle on several tasks at long contexts, underscoring both the usefulness of LC-Eval for benchmarking and the ongoing need to improve multilingual long-context capabilities, particularly in Arabic.

Abstract

Recent advancements in Large Language Models (LLMs) have demonstrated sophisticated capabilities, including the ability to process and comprehend extended contexts. These emergent capabilities necessitate rigorous evaluation methods to effectively assess their performance in long-context understanding. In this paper, we present \textbf{LC-Eval}, a bilingual, multi-task evaluation benchmark designed to evaluate long-context understanding in English and Arabic, targeting context lengths ranging from 4k to over 128k tokens. LC-Eval introduces four novel and challenging tasks: multi-document question answering, bilingual question answering, claim verification within a paragraph, and multiple-choice questions based on long contexts. These tasks are designed to assess LLMs' abilities in deep reasoning, document comprehension, information tracing, and bilingual information extraction and understanding. The benchmark includes datasets in both Arabic and English for each task, allowing for a comparative analysis of their performance across different text genres. Evaluations were conducted on both open-weight and closed LLMs, with results indicating that LC-Eval presents significant challenges. Even high-performing models, such as GPT-4o, struggled with certain tasks, highlighting the complexity and rigor of the benchmark.

LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding

TL;DR

LC-Eval introduces a bilingual multi-task benchmark to assess long-context understanding in English and Arabic, covering 4K to over 128K tokens across four tasks: multi-document QA, bilingual QA, claim verification, and MCQ QA. The authors augment the dataset with an entity-relationship–based evaluation that uses LLMs as semantic judges, and validate all data with multiple human annotators. The dataset comprises 7,903 samples drawn from diverse public sources and includes careful curation to stress deep reasoning, information tracing, and cross-lingual extraction. Experimental results show that even strong models like GPT-4o struggle on several tasks at long contexts, underscoring both the usefulness of LC-Eval for benchmarking and the ongoing need to improve multilingual long-context capabilities, particularly in Arabic.

Abstract

Recent advancements in Large Language Models (LLMs) have demonstrated sophisticated capabilities, including the ability to process and comprehend extended contexts. These emergent capabilities necessitate rigorous evaluation methods to effectively assess their performance in long-context understanding. In this paper, we present \textbf{LC-Eval}, a bilingual, multi-task evaluation benchmark designed to evaluate long-context understanding in English and Arabic, targeting context lengths ranging from 4k to over 128k tokens. LC-Eval introduces four novel and challenging tasks: multi-document question answering, bilingual question answering, claim verification within a paragraph, and multiple-choice questions based on long contexts. These tasks are designed to assess LLMs' abilities in deep reasoning, document comprehension, information tracing, and bilingual information extraction and understanding. The benchmark includes datasets in both Arabic and English for each task, allowing for a comparative analysis of their performance across different text genres. Evaluations were conducted on both open-weight and closed LLMs, with results indicating that LC-Eval presents significant challenges. Even high-performing models, such as GPT-4o, struggled with certain tasks, highlighting the complexity and rigor of the benchmark.
Paper Structure (67 sections, 1 equation, 1 figure, 15 tables)