Table of Contents
Fetching ...

Data-intrinsic approximation in metric spaces

Jürgen Dölz, Michael Multerer

TL;DR

The paper addresses the challenge of approximating labeled data mapped on finite metric spaces without imposing strong model assumptions. It introduces the discrete modulus of continuity as a data-intrinsic regularity measure and develops both deterministic and probabilistic consistency results, linking regularity to covering numbers and random ball covers. It further provides computational strategies, including efficient nearest-neighbor and set-cover algorithms, and extends the framework with piecewise-constant interpolation and multilevel Monte Carlo for empirical data. Numerical experiments on synthetic and real data validate the theory and demonstrate the practical viability of the data-centric approach for high-dimensional, irregular data on general metric spaces.

Abstract

Analysis and processing of data is a vital part of our modern society and requires vast amounts of computational resources. To reduce the computational burden, compressing and approximating data has become a central topic. We consider the approximation of labeled data samples, mathematically described as site-to-value maps between finite metric spaces. Within this setting, we identify the discrete modulus of continuity as an effective data-intrinsic quantity to measure regularity of site-to-value maps without imposing further structural assumptions. We investigate the consistency of the discrete modulus of continuity in the infinite data limit and propose an algorithm for its efficient computation. Building on these results, we present a sample based approximation theory for labeled data. For data subject to statistical uncertainty we consider multilevel approximation spaces and a variant of the multilevel Monte Carlo method to compute statistical quantities of interest. Our considerations connect approximation theory for labeled data in metric spaces to the covering problem for (random) balls on the one hand and the efficient evaluation of the discrete modulus of continuity to combinatorial optimization on the other hand. We provide extensive numerical studies to illustrate the feasibility of the approach and to validate our theoretical results.

Data-intrinsic approximation in metric spaces

TL;DR

The paper addresses the challenge of approximating labeled data mapped on finite metric spaces without imposing strong model assumptions. It introduces the discrete modulus of continuity as a data-intrinsic regularity measure and develops both deterministic and probabilistic consistency results, linking regularity to covering numbers and random ball covers. It further provides computational strategies, including efficient nearest-neighbor and set-cover algorithms, and extends the framework with piecewise-constant interpolation and multilevel Monte Carlo for empirical data. Numerical experiments on synthetic and real data validate the theory and demonstrate the practical viability of the data-centric approach for high-dimensional, irregular data on general metric spaces.

Abstract

Analysis and processing of data is a vital part of our modern society and requires vast amounts of computational resources. To reduce the computational burden, compressing and approximating data has become a central topic. We consider the approximation of labeled data samples, mathematically described as site-to-value maps between finite metric spaces. Within this setting, we identify the discrete modulus of continuity as an effective data-intrinsic quantity to measure regularity of site-to-value maps without imposing further structural assumptions. We investigate the consistency of the discrete modulus of continuity in the infinite data limit and propose an algorithm for its efficient computation. Building on these results, we present a sample based approximation theory for labeled data. For data subject to statistical uncertainty we consider multilevel approximation spaces and a variant of the multilevel Monte Carlo method to compute statistical quantities of interest. Our considerations connect approximation theory for labeled data in metric spaces to the covering problem for (random) balls on the one hand and the efficient evaluation of the discrete modulus of continuity to combinatorial optimization on the other hand. We provide extensive numerical studies to illustrate the feasibility of the approach and to validate our theoretical results.
Paper Structure (41 sections, 32 theorems, 160 equations, 15 figures, 1 table, 4 algorithms)

This paper contains 41 sections, 32 theorems, 160 equations, 15 figures, 1 table, 4 algorithms.

Key Result

Lemma 2.4

Let $f\colon \mathcal{X}\to \mathcal{Y}$ be uniformly continuous. Then, the modulus of continuity

Figures (15)

  • Figure 1: Illustration of two continuous functions $f_N,g_N\colon\{1,\ldots,6\}\to \mathbb{N}$ (left) and their corresponding discrete, optimal, right-continuous moduli of continuity as a measure of their smoothness (right).
  • Figure 2: Visualization of the situation in \ref{['eq:nearestmodneighbor']}.
  • Figure 3: Samples of the discrete modulus of continuity of the Wiener process (left). Average modulus of continuity and one standard deviation envelope (right). Dashed lines indicate the approximation \ref{['eq:wienermodcont']} for small $t$.
  • Figure 4: Continuous but not Hölder continuous function from \ref{['eq:nothoelder']} on $\mathbb{S}^2$ (left) and corresponding discrete modulus of continuity (right). The dashed line indicates the exact modulus of continuity.
  • Figure 5: Time series of the real-world weather data (left) and its discrete modulus of continuity (right).
  • ...and 10 more figures

Theorems & Definitions (75)

  • Definition 2.1
  • Definition 2.2
  • Remark 2.3
  • Lemma 2.4: Properties of the modulus of continuity
  • Remark 2.5
  • Remark 2.6
  • Lemma 2.7: Monotonicity of discrete modulus of continuity
  • proof
  • Lemma 2.8: Monotonicity of the discrete seminorm
  • proof
  • ...and 65 more