Information Theory in Open-world Machine Learning Foundations, Frameworks, and Future Direction
Lin Wang
TL;DR
This paper surveys information-theoretic approaches to Open-world Machine Learning (OWML), arguing that core quantities like entropy $H$, mutual information $I$, and KL divergence $D_{\mathrm{KL}}$ provide a unifying language for uncertainty, knowledge transfer, and adaptation in open, nonstationary environments. It formalizes OWML as an information-flow problem, extending the Information Bottleneck to include unknown/nonstationary components, and presents objective formulations that balance compression, retention of known knowledge, and suppression of unknown information. The review covers applications to Open-set Recognition, Novelty Discovery, and Continual Learning, and connects these to provable learning via PAC-Bayes bounds, open-space risk, and causal information flow. Finally, it outlines open problems and directions toward a theory of Provable Information Dynamics, emphasizing dynamic bounds, multimodal information, causality integration, and self-adaptive agents, with the aim of enabling provable, trustworthy open-world AI.
Abstract
Open world Machine Learning (OWML) aims to develop intelligent systems capable of recognizing known categories, rejecting unknown samples, and continually learning from novel information. Despite significant progress in open set recognition, novelty detection, and continual learning, the field still lacks a unified theoretical foundation that can quantify uncertainty, characterize information transfer, and explain learning adaptability in dynamic, nonstationary environments. This paper presents a comprehensive review of information theoretic approaches in open world machine learning, emphasizing how core concepts such as entropy, mutual information, and Kullback Leibler divergence provide a mathematical language for describing knowledge acquisition, uncertainty suppression, and risk control under open world conditions. We synthesize recent studies into three major research axes: information theoretic open set recognition enabling safe rejection of unknowns, information driven novelty discovery guiding new concept formation, and information retentive continual learning ensuring stable long term adaptation. Furthermore, we discuss theoretical connections between information theory and provable learning frameworks, including PAC Bayes bounds, open-space risk theory, and causal information flow, to establish a pathway toward provable and trustworthy open world intelligence. Finally, the review identifies key open problems and future research directions, such as the quantification of information risk, development of dynamic mutual information bounds, multimodal information fusion, and integration of information theory with causal reasoning and world model learning.
