Table of Contents
Fetching ...

FINDER: Feature Inference on Noisy Datasets using Eigenspace Residuals

Trajan Murphy, Akshunna S. Dogra, Hanfeng Gu, Caleb Meredith, Mark Kon, Julio Enrique Castrillion-Candas

TL;DR

FINDER tackles classification in data-scarce and noisy regimes by introducing stochastic features rooted in a generalized Kosambi-Karhunen–Loève expansion, embedding empirical datasets into a Hilbert space where class separation can be analyzed spectrally. The framework centers on constructing residual eigenspaces $\mathcal{H}_{\text{res}}$ to achieve distinct spectral profiles for different classes, with variants MLS and ACA-S/ACA-L providing practical implementation strategies. Empirical results in Alzheimer's disease proteomics and remote sensing deforestation demonstrate state-of-the-art improvements in AUC/accuracy and computational efficiency, especially under data-poor conditions, while highlighting the method’s robustness and potential integration with simpler classifiers like SVMs or HMMs. Limitations include sensitivity to truncation choices, suboptimal probabilistic bounds, and the binary nature of the current formulation, guiding avenues for future work and broader applicability in multi-class settings.

Abstract

''Noisy'' datasets (regimes with low signal to noise ratios, small sample sizes, faulty data collection, etc) remain a key research frontier for classification methods with both theoretical and practical implications. We introduce FINDER, a rigorous framework for analyzing generic classification problems, with tailored algorithms for noisy datasets. FINDER incorporates fundamental stochastic analysis ideas into the feature learning and inference stages to optimally account for the randomness inherent to all empirical datasets. We construct ''stochastic features'' by first viewing empirical datasets as realizations from an underlying random field (without assumptions on its exact distribution) and then mapping them to appropriate Hilbert spaces. The Kosambi-Karhunen-Loéve expansion (KLE) breaks these stochastic features into computable irreducible components, which allow classification over noisy datasets via an eigen-decomposition: data from different classes resides in distinct regions, identified by analyzing the spectrum of the associated operators. We validate FINDER on several challenging, data-deficient scientific domains, producing state of the art breakthroughs in: (i) Alzheimer's Disease stage classification, (ii) Remote sensing detection of deforestation. We end with a discussion on when FINDER is expected to outperform existing methods, its failure modes, and other limitations.

FINDER: Feature Inference on Noisy Datasets using Eigenspace Residuals

TL;DR

FINDER tackles classification in data-scarce and noisy regimes by introducing stochastic features rooted in a generalized Kosambi-Karhunen–Loève expansion, embedding empirical datasets into a Hilbert space where class separation can be analyzed spectrally. The framework centers on constructing residual eigenspaces to achieve distinct spectral profiles for different classes, with variants MLS and ACA-S/ACA-L providing practical implementation strategies. Empirical results in Alzheimer's disease proteomics and remote sensing deforestation demonstrate state-of-the-art improvements in AUC/accuracy and computational efficiency, especially under data-poor conditions, while highlighting the method’s robustness and potential integration with simpler classifiers like SVMs or HMMs. Limitations include sensitivity to truncation choices, suboptimal probabilistic bounds, and the binary nature of the current formulation, guiding avenues for future work and broader applicability in multi-class settings.

Abstract

''Noisy'' datasets (regimes with low signal to noise ratios, small sample sizes, faulty data collection, etc) remain a key research frontier for classification methods with both theoretical and practical implications. We introduce FINDER, a rigorous framework for analyzing generic classification problems, with tailored algorithms for noisy datasets. FINDER incorporates fundamental stochastic analysis ideas into the feature learning and inference stages to optimally account for the randomness inherent to all empirical datasets. We construct ''stochastic features'' by first viewing empirical datasets as realizations from an underlying random field (without assumptions on its exact distribution) and then mapping them to appropriate Hilbert spaces. The Kosambi-Karhunen-Loéve expansion (KLE) breaks these stochastic features into computable irreducible components, which allow classification over noisy datasets via an eigen-decomposition: data from different classes resides in distinct regions, identified by analyzing the spectrum of the associated operators. We validate FINDER on several challenging, data-deficient scientific domains, producing state of the art breakthroughs in: (i) Alzheimer's Disease stage classification, (ii) Remote sensing detection of deforestation. We end with a discussion on when FINDER is expected to outperform existing methods, its failure modes, and other limitations.
Paper Structure (31 sections, 7 theorems, 27 equations, 11 figures, 8 tables)

This paper contains 31 sections, 7 theorems, 27 equations, 11 figures, 8 tables.

Key Result

Theorem 2.1

Let $v \in L^2(\Omega, \mathcal{H})$ be Bochner measurable. Then, there exists $R \in \mathbb{N} \cup \aleph_0$ such that where $\left \lbrace Y_r \right \rbrace_{r=1}^R, \left \lbrace \phi_r \right \rbrace_{r=1}^R$ are orthonormal sets in $L^2(\Omega), \mathcal{H}$ respectively, with $\mathbb{E}[Y_r] = 0$ for all $r$.

Figures (11)

  • Figure 1: A schematic for FINDER and a visual perspective on classification as a multi-stage process.
  • Figure 2: AUC obtained across all three methods for each of the three ADNI cohorts. Within each method (MLS, ACA, benchmark), the regime with the highest overall AUC obtained across all tested values of $M_\text{res}$ is reported. AUC can improve significantly from the benchmark level when both FINDER methods are employed with an RBF separating boundary. While the MLS method remains robust with respect to the choice to pre-balance or not pre-balanced the data, the ACA method is highly dependent on this choice. For the AD vs. CN cohort, the Unbalanced regime within ACA performs consistently better than benchmark and MLS. However, for the CN vs. LMCI cohort, the Balanced regime within ACA performs consistently better than benchmark and MLS. The data also demonstrate that the performance of FINDER is sensitive to the choice of $M_\text{res}$. The overall trend appears to be that larger $M_\text{res}$ achieve higher AUCs, though too large $M_\text{res}$ can diminish AUCs.
  • Figure 3: (a) Test region in the Amazon forest and validation samples. 1000 samples of the validation regions are selected. The region is formed by $9,219 \times 9,180$ pixels, each pixel a $10 m \times 10m$ patch of land, representing the Enhanced Vegetation Index. The colored areas indicate the detection of deforestation with the Hybrid FINDER+HMM method by December 31 2022. (b) Overall metric accuracy with Hybrid and Optical only vs the number of available optical Sentinel-2 days.
  • Figure 4: The division of the dataset $\mathcal{D}$ into training (which includes SVM separating hypersurface and covariance operator estimation) and validation (testing) subsets for one round of LPOCV.
  • Figure 5: Accuracy obtained across all three methods for each of the three ADNI cohorts. Within each method (benchmark, MLS, ACA), we display the regime which obtains the maximum accuracy among the values of $M_\text{res}$ tested. We observe that the ACA-S regime either outperforms or matches both the MLS method and the benchmark learners across all three cohorts. However, the performance of the ACA-S regime can be sensitive to the choice of $M_\text{res}$, as exemplified in the AD vs. LMCI and CN vs. LMCI cohorts. Values of $M_\text{res}$ which are too high or too low can prevent the ACA-S regime from achieving maximum accuracy, and can even reduce the accuracy below benchmark level, as exhibited by the AD vs. LMCI cohort. The data demonstrate a significant improvement on the benchmark performance, prompting an ad-hoc analysis of the ratio of the FINDER and benchmark error rates as tabulated in Figure \ref{['ADNI accuracy refinement']}.
  • ...and 6 more figures

Theorems & Definitions (23)

  • Definition 2.1
  • Theorem 2.1
  • Lemma 2.1
  • Lemma 2.2
  • Lemma 2.3
  • Remark 2.1
  • Remark 3.1
  • Definition A.1: Simple functions
  • Definition A.2: Bochner-measurable
  • Remark A.1
  • ...and 13 more