CDF Transform-and-Shift: An effective way to deal with datasets of inhomogeneous cluster densities

Ye Zhu; Kai Ming Ting; Mark Carman; Maia Angelova

CDF Transform-and-Shift: An effective way to deal with datasets of inhomogeneous cluster densities

Ye Zhu, Kai Ming Ting, Mark Carman, Maia Angelova

TL;DR

The paper tackles the problem of biased clustering and anomaly detection in datasets with inhomogeneous cluster densities. It introduces CDF Transform-and-Shift (CDF-TS), a multi-dimensional CDF-based preprocessing that homogenises local densities while preserving cluster structure, enabling existing algorithms to operate under their implicit assumptions without modification. Through extensive experiments on clustering and $k$NN anomaly detection, CDF-TS consistently improves performance over state-of-the-art remedies like ReScale and DScale, and is compatible with multiple density estimators. The method provides a practical, general approach to mitigating density bias and can extend to other density-based techniques beyond those tested.

Abstract

The problem of inhomogeneous cluster densities has been a long-standing issue for distance-based and density-based algorithms in clustering and anomaly detection. These algorithms implicitly assume that all clusters have approximately the same density. As a result, they often exhibit a bias towards dense clusters in the presence of sparse clusters. Many remedies have been suggested; yet, we show that they are partial solutions which do not address the issue satisfactorily. To match the implicit assumption, we propose to transform a given dataset such that the transformed clusters have approximately the same density while all regions of locally low density become globally low density -- homogenising cluster density while preserving the cluster structure of the dataset. We show that this can be achieved by using a new multi-dimensional Cumulative Distribution Function in a transform-and-shift method. The method can be applied to every dataset, before the dataset is used in many existing algorithms to match their implicit assumption without algorithmic modification. We show that the proposed method performs better than existing remedies.

CDF Transform-and-Shift: An effective way to deal with datasets of inhomogeneous cluster densities

TL;DR

Abstract

CDF Transform-and-Shift: An effective way to deal with datasets of inhomogeneous cluster densities

TL;DR

Abstract

Paper Structure

Table of Contents

Key Result

Figures (12)

Theorems & Definitions (5)