Table of Contents
Fetching ...

Concentration and excess risk bounds for imbalanced classification with synthetic oversampling

Touqeer Ahmad, Mohammadreza M. Kalan, François Portier, Gilles Stupfler

TL;DR

A theoretical framework to analyze the behavior of SMOTE and related methods when classifiers are trained on synthetic data is developed and a nonparametric excess risk guarantee is provided for kernel-based classifiers trained using such synthetic data.

Abstract

Synthetic oversampling of minority examples using SMOTE and its variants is a leading strategy for addressing imbalanced classification problems. Despite the success of this approach in practice, its theoretical foundations remain underexplored. We develop a theoretical framework to analyze the behavior of SMOTE and related methods when classifiers are trained on synthetic data. We first derive a uniform concentration bound on the discrepancy between the empirical risk over synthetic minority samples and the population risk on the true minority distribution. We then provide a nonparametric excess risk guarantee for kernel-based classifiers trained using such synthetic data. These results lead to practical guidelines for better parameter tuning of both SMOTE and the downstream learning algorithm. Numerical experiments are provided to illustrate and support the theoretical findings

Concentration and excess risk bounds for imbalanced classification with synthetic oversampling

TL;DR

A theoretical framework to analyze the behavior of SMOTE and related methods when classifiers are trained on synthetic data is developed and a nonparametric excess risk guarantee is provided for kernel-based classifiers trained using such synthetic data.

Abstract

Synthetic oversampling of minority examples using SMOTE and its variants is a leading strategy for addressing imbalanced classification problems. Despite the success of this approach in practice, its theoretical foundations remain underexplored. We develop a theoretical framework to analyze the behavior of SMOTE and related methods when classifiers are trained on synthetic data. We first derive a uniform concentration bound on the discrepancy between the empirical risk over synthetic minority samples and the population risk on the true minority distribution. We then provide a nonparametric excess risk guarantee for kernel-based classifiers trained using such synthetic data. These results lead to practical guidelines for better parameter tuning of both SMOTE and the downstream learning algorithm. Numerical experiments are provided to illustrate and support the theoretical findings
Paper Structure (26 sections, 19 theorems, 131 equations, 12 figures, 2 algorithms)

This paper contains 26 sections, 19 theorems, 131 equations, 12 figures, 2 algorithms.

Key Result

Theorem 1

Suppose that Assumptions assump_regularity, function_class and assump:rad-free hold. Let $\delta \in (0,1/5)$ and $m\geq 1$. If $n_1>1$, let $k \in \{1,\ldots, n_1-1\}$ and let $\{X_{1i}^{*}\}_{1\leq i\leq m}$ be $m$ i.i.d. samples generated by the Smote algorithm smote-equation. Then, with probabil where

Figures (12)

  • Figure 1: Average AM-risk of KNN, KS, and LR classifiers on balanced data over 50 replications. Left: using Smote and Smote($k$) with $k \in (7, 65)$. Right: using Kdeo and Kdeo$(H)$, with $H=cH_1$ and $c$ ranging in $(1/20, 3)$ where $H_1$ follows from Scott’s rule.
  • Figure 2: Average AM-risk across different data imbalance regimes for the KS (described in Section \ref{['sec:main:excessrisk']}) and $K$NN classification rules computed over 50 replications.
  • Figure 3: Average AM-risk across different data imbalance regimes for the $K$NN methods described in Section \ref{['Methods+variants']}, computed over 50 replications.
  • Figure 4: AM-risk corresponding to different rebalancing methods and datasets when using the $K$NN classifier.
  • Figure 5: Average AM-risk of KNN, KS, and LR classifiers on balanced data over 50 replications. Left: using Smote and Smote($k$) with $k \in (7, 65)$. Right: using Kdeo and Kdeo$(H)$, with $H=cH_1$ and $c$ ranging in $(1/20, 3)$ where $H_1$ follows from Scott’s rule.
  • ...and 7 more figures

Theorems & Definitions (33)

  • Definition 1: Rademacher complexity
  • Theorem 1
  • Corollary 2
  • Theorem 3
  • Corollary 4
  • Remark 1
  • Remark 2
  • Remark 3
  • Theorem 5
  • Theorem 6
  • ...and 23 more