Exploring Complexity Changes in Diseased ECG Signals for Enhanced Classification
Camilo Quiceno Quintero, Sandip Varkey George
TL;DR
This work investigates how nonlinear ECG complexity metrics reflect cardiac pathology and whether they can improve automated disease detection. Using the PTB-XL dataset, the authors compute a broad set of single-lead and cross-lead nonlinear measures (including $HD$, $ApEn$, $PermEn$, $LZC$, $MSE$, RP-derived features, $\rho_S$, and $MI$) and evaluate their utility in binary and five-class classifications with multiple ML models. They demonstrate widespread differences between healthy and diseased ECGs ($p<.001$) and gains in classification performance when incorporating complexity features, particularly cross-lead metrics, with ANN achieving AUCs up to $0.91$ in binary tasks and around $0.70$ accuracy in five-class tasks. The findings highlight the sensitivity of ECG dynamics to pathology and suggest that complexity-based features can enhance diagnostic tools and disease stratification in clinical settings.
Abstract
The complex dynamics of the heart are reflected in its electrical activity, captured through electrocardiograms (ECGs). In this study we use nonlinear time series analysis to understand how ECG complexity varies with cardiac pathology. Using the large PTB-XL dataset, we extracted nonlinear measures from lead II ECGs, and cross-channel metrics (leads II, V2, AVL) using Spearman correlations and mutual information. Significant differences between diseased and healthy individuals were found in almost all measures between healthy and diseased classes, and between 5 diagnostic superclasses ($p<.001$). Moreover, incorporating these complexity quantifiers into machine learning models substantially improved classification accuracy measured using area under the ROC curve (AUC) from 0.86 (baseline) to 0.87 (nonlinear measures) and 0.90 (including cross-time series metrics).
