ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains

Yein Park; Chanwoong Yoon; Jungwoo Park; Donghyeon Lee; Minbyul Jeong; Jaewoo Kang

ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains

Yein Park, Chanwoong Yoon, Jungwoo Park, Donghyeon Lee, Minbyul Jeong, Jaewoo Kang

TL;DR

This work tackles the challenge of evaluating and improving temporal knowledge in large language models (LLMs) by introducing ChroKnowBench, a benchmark that tracks time-variant and time-invariant knowledge across multiple domains with dynamic and static states. It couples this benchmark with ChroKnowledge, a sampling-based framework, and ChroKnowPrompt, an in-depth chronological prompting method, to non-parametrically elicit and refine temporal knowledge. Key findings show domain characteristics strongly influence temporal knowledge representation, with improved recall for unchanged objects and limited gains for dynamic changes when prompted alone. The proposed approach highlights the importance of temporal context for up-to-date reasoning and suggests future work combining non-parametric prompting with parametric updates to better capture evolving facts in practice.

Abstract

Large language models (LLMs) have brought significant changes to many aspects of our lives. However, assessing and ensuring their chronological knowledge remains challenging. Existing approaches fall short in addressing the temporal adaptability of knowledge, often relying on a fixed time-point view. To overcome this, we introduce ChroKnowBench, a benchmark dataset designed to evaluate chronologically accumulated knowledge across three key aspects: multiple domains, time dependency, temporal state. Our benchmark distinguishes between knowledge that evolves (e.g., personal history, scientific discoveries, amended laws) and knowledge that remain constant (e.g., mathematical truths, commonsense facts). Building on this benchmark, we present ChroKnowledge (Chronological Categorization of Knowledge), a novel sampling-based framework for evaluating LLMs' non-parametric chronological knowledge. Our evaluation led to the following observations: (1) The ability of eliciting temporal knowledge varies depending on the data format that model was trained on. (2) LLMs partially recall knowledge or show a cut-off at temporal boundaries rather than recalling all aspects of knowledge correctly. Thus, we apply our ChroKnowPrompt, an in-depth prompting to elicit chronological knowledge by traversing step-by-step through the surrounding time spans. We observe that it successfully recalls objects across both open-source and proprietary LLMs, demonstrating versatility, though it faces challenges with dynamic datasets and unstructured formats.

ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains

TL;DR

Abstract

ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (17)