Table of Contents
Fetching ...

Data Understanding Survey: Pursuing Improved Dataset Characterization Via Tensor-based Methods

Matthew D. Merris, Tim Andersen

TL;DR

Data understanding is essential for robust, explainable AI. The paper surveys conventional dataset characterization methods and argues that tensor-based data analysis can reveal higher-order data characteristics that improve explainability. It reviews tensor theory, decompositions such as CP and Tucker, generalized variants like GCP, and tensor-based methods for clustering and model-based analysis, illustrating how tensorization and implicit formulations can circumvent computational bottlenecks. The authors present a framework mapping tensor methods to canonical DC categories and propose directions where tensorization and tensor moments can advance DC in big data settings.

Abstract

In the evolving domains of Machine Learning and Data Analytics, existing dataset characterization methods such as statistical, structural, and model-based analyses often fail to deliver the deep understanding and insights essential for innovation and explainability. This work surveys the current state-of-the-art conventional data analytic techniques and examines their limitations, and discusses a variety of tensor-based methods and how these may provide a more robust alternative to traditional statistical, structural, and model-based dataset characterization techniques. Through examples, we illustrate how tensor methods unveil nuanced data characteristics, offering enhanced interpretability and actionable intelligence. We advocate for the adoption of tensor-based characterization, promising a leap forward in understanding complex datasets and paving the way for intelligent, explainable data-driven discoveries.

Data Understanding Survey: Pursuing Improved Dataset Characterization Via Tensor-based Methods

TL;DR

Data understanding is essential for robust, explainable AI. The paper surveys conventional dataset characterization methods and argues that tensor-based data analysis can reveal higher-order data characteristics that improve explainability. It reviews tensor theory, decompositions such as CP and Tucker, generalized variants like GCP, and tensor-based methods for clustering and model-based analysis, illustrating how tensorization and implicit formulations can circumvent computational bottlenecks. The authors present a framework mapping tensor methods to canonical DC categories and propose directions where tensorization and tensor moments can advance DC in big data settings.

Abstract

In the evolving domains of Machine Learning and Data Analytics, existing dataset characterization methods such as statistical, structural, and model-based analyses often fail to deliver the deep understanding and insights essential for innovation and explainability. This work surveys the current state-of-the-art conventional data analytic techniques and examines their limitations, and discusses a variety of tensor-based methods and how these may provide a more robust alternative to traditional statistical, structural, and model-based dataset characterization techniques. Through examples, we illustrate how tensor methods unveil nuanced data characteristics, offering enhanced interpretability and actionable intelligence. We advocate for the adoption of tensor-based characterization, promising a leap forward in understanding complex datasets and paving the way for intelligent, explainable data-driven discoveries.
Paper Structure (48 sections, 61 equations, 8 figures, 2 tables)

This paper contains 48 sections, 61 equations, 8 figures, 2 tables.

Figures (8)

  • Figure 1: Sampling of the Datasaurus Dozen. Each dataset share the same summary statistics to two decimal places.
  • Figure 2: Workflow for comparing standard statistics between two datasets robnik-sikonjaDatasetComparisonWorkflows2018
  • Figure 3: Workflow for comparing clustering similarity between two datasets robnik-sikonjaDatasetComparisonWorkflows2018
  • Figure 4: Workflow for comparing classification performance between two datasets robnik-sikonjaDatasetComparisonWorkflows2018
  • Figure 5: Forward (top) and Backward (bottom) Mode-1 matricization of a tensor.
  • ...and 3 more figures