Data Understanding Survey: Pursuing Improved Dataset Characterization Via Tensor-based Methods
Matthew D. Merris, Tim Andersen
TL;DR
Data understanding is essential for robust, explainable AI. The paper surveys conventional dataset characterization methods and argues that tensor-based data analysis can reveal higher-order data characteristics that improve explainability. It reviews tensor theory, decompositions such as CP and Tucker, generalized variants like GCP, and tensor-based methods for clustering and model-based analysis, illustrating how tensorization and implicit formulations can circumvent computational bottlenecks. The authors present a framework mapping tensor methods to canonical DC categories and propose directions where tensorization and tensor moments can advance DC in big data settings.
Abstract
In the evolving domains of Machine Learning and Data Analytics, existing dataset characterization methods such as statistical, structural, and model-based analyses often fail to deliver the deep understanding and insights essential for innovation and explainability. This work surveys the current state-of-the-art conventional data analytic techniques and examines their limitations, and discusses a variety of tensor-based methods and how these may provide a more robust alternative to traditional statistical, structural, and model-based dataset characterization techniques. Through examples, we illustrate how tensor methods unveil nuanced data characteristics, offering enhanced interpretability and actionable intelligence. We advocate for the adoption of tensor-based characterization, promising a leap forward in understanding complex datasets and paving the way for intelligent, explainable data-driven discoveries.
