Table of Contents
Fetching ...

Local regression on path spaces with signature metrics

Christian Bayer, Davit Gogolashvili, Luca Pelizzari

TL;DR

The paper develops a local regression framework for path-valued data by embedding paths via the signature transform and using a signature-induced semi-metric within a Nadaraya–Watson estimator. It provides finite-sample convergence guarantees, with error rates governed by small-ball probabilities; the authors show that truncating the signature yields Euclidean-like rates with effective dimension $\nu(N)$, while the full signature space can incur slower, exponential-type rates. They extend the theory to rough differential equations, where the Itô–Lyons map is Hölder continuous in signature distance, and address robustness through a robust signature variant. Empirically, the method improves learning of SDE solution maps and achieves competitive time-series classification results with favorable computational efficiency compared to kernel-based methods and DTW baselines. The framework offers a unified, scalable approach for nonparametric regression and classification on infinite-dimensional path spaces with solid theoretical guarantees and practical impact for sequential data analysis.

Abstract

We study nonparametric regression and classification for path-valued data. We introduce a functional Nadaraya-Watson estimator that combines the signature transform from rough path theory with local kernel regression. The signature transform provides a principled way to encode sequential data through iterated integrals, enabling direct comparison of paths in a natural metric space. Our approach leverages signature-induced distances within the classical kernel regression framework, achieving computational efficiency while avoiding the scalability bottlenecks of large-scale kernel matrix operations. We establish finite-sample convergence bounds demonstrating favorable statistical properties of signature-based distances compared to traditional metrics in infinite-dimensional settings. We propose robust signature variants that provide stability against outliers, enhancing practical performance. Applications to both synthetic and real-world data - including stochastic differential equation learning and time series classification - demonstrate competitive accuracy while offering significant computational advantages over existing methods.

Local regression on path spaces with signature metrics

TL;DR

The paper develops a local regression framework for path-valued data by embedding paths via the signature transform and using a signature-induced semi-metric within a Nadaraya–Watson estimator. It provides finite-sample convergence guarantees, with error rates governed by small-ball probabilities; the authors show that truncating the signature yields Euclidean-like rates with effective dimension , while the full signature space can incur slower, exponential-type rates. They extend the theory to rough differential equations, where the Itô–Lyons map is Hölder continuous in signature distance, and address robustness through a robust signature variant. Empirically, the method improves learning of SDE solution maps and achieves competitive time-series classification results with favorable computational efficiency compared to kernel-based methods and DTW baselines. The framework offers a unified, scalable approach for nonparametric regression and classification on infinite-dimensional path spaces with solid theoretical guarantees and practical impact for sequential data analysis.

Abstract

We study nonparametric regression and classification for path-valued data. We introduce a functional Nadaraya-Watson estimator that combines the signature transform from rough path theory with local kernel regression. The signature transform provides a principled way to encode sequential data through iterated integrals, enabling direct comparison of paths in a natural metric space. Our approach leverages signature-induced distances within the classical kernel regression framework, achieving computational efficiency while avoiding the scalability bottlenecks of large-scale kernel matrix operations. We establish finite-sample convergence bounds demonstrating favorable statistical properties of signature-based distances compared to traditional metrics in infinite-dimensional settings. We propose robust signature variants that provide stability against outliers, enhancing practical performance. Applications to both synthetic and real-world data - including stochastic differential equation learning and time series classification - demonstrate competitive accuracy while offering significant computational advantages over existing methods.
Paper Structure (20 sections, 12 theorems, 116 equations, 1 figure, 2 tables)

This paper contains 20 sections, 12 theorems, 116 equations, 1 figure, 2 tables.

Key Result

Theorem 4

Let $Y \in [-R,R]$ and let $F \in \mathcal{F}_{\beta}$ with smoothness parameter $\beta \in (0,1]$. Consider the estimator $\widehat{F}$ defined in eq:NW_estimator_regression, and assume that the kernel $K$ is compactly supported and satisfies for some constants $0 < b \leq B < \infty$. For any $\delta \in (0,1)$ and $M$ satisfying with probability at least $1-\delta$ the following bound holds

Figures (1)

  • Figure 1: Scatter plots of the testing data $\{(Y^{(m)}, \widehat{Y}^{(m)}):m \in I_{te} \}$, using signature and supremum metrics in \ref{['eq:estimator_SDE']}. At truncation level $N=4$, the point ${\color{orange} \bullet}$ near $(1.2,0)$ illustrates an outlier of the basic method Sig; see the discussion in Section \ref{['sec:robustification']}.

Theorems & Definitions (28)

  • Definition 1
  • Remark 2
  • Definition 3
  • Theorem 4
  • Remark 5
  • Remark 6
  • Lemma 7
  • Example 1
  • Proposition 9
  • Remark 10
  • ...and 18 more