Area under the ROC Curve has the Most Consistent Evaluation for Binary Classification

Jing Li

Area under the ROC Curve has the Most Consistent Evaluation for Binary Classification

Jing Li

TL;DR

The paper addresses how binary classification evaluation metrics vary with data prevalence and demonstrates that Area Under the ROC Curve (AUC) offers the most consistent evaluation across prevalence changes. Using statistical simulations across 156 data scenarios, 18 metrics, 5 common models, and a random baseline, it shows that metrics influenced by prevalence exhibit higher variance in both model evaluation and rankings, while AUC remains largely unaffected by prevalence due to its threshold-agnostic nature. A threshold-analysis reveals that considering more decision thresholds further reduces variance, bolstering the case for AUC as a robust summary metric. The findings have practical implications for model evaluation and selection, particularly in settings with changing class distributions, and point to AUC as a reliable default when cross-scenario consistency is desired.

Abstract

The proper use of model evaluation metrics is important for model evaluation and model selection in binary classification tasks. This study investigates how consistent different metrics are at evaluating models across data of different prevalence while the relationships between different variables and the sample size are kept constant. Analyzing 156 data scenarios, 18 model evaluation metrics and five commonly used machine learning models as well as a naive random guess model, I find that evaluation metrics that are less influenced by prevalence offer more consistent evaluation of individual models and more consistent ranking of a set of models. In particular, Area Under the ROC Curve (AUC) which takes all decision thresholds into account when evaluating models has the smallest variance in evaluating individual models and smallest variance in ranking of a set of models. A close threshold analysis using all possible thresholds for all metrics further supports the hypothesis that considering all decision thresholds helps reduce the variance in model evaluation with respect to prevalence change in data. The results have significant implications for model evaluation and model selection in binary classification tasks.

Area under the ROC Curve has the Most Consistent Evaluation for Binary Classification

TL;DR

Abstract

Paper Structure (17 sections, 13 equations, 14 figures, 3 tables)

This paper contains 17 sections, 13 equations, 14 figures, 3 tables.

Introduction
Methods
Background and Basic Notations
Alternative Derivation of Confusion Matrix Metrics
Case Study: Crime Recidivism
Simulation Study
Results
Simulation Results: Evaluation Metrics
Simulation Results: Comparisons of Models
Simulation Results: Ranking of Different Models
Threshold Analysis
Discussions and Conclusion
Discussions
Conclusion
Prediction Results for Crime Recidivism
...and 2 more sections

Figures (14)

Figure 1: Model Evaluation by different Metrics at different prevalence
Figure 2: Model Comparison by different metrics
Figure 3: Model Rankings by different metrics
Figure 4: Model Evaluation with Different thresholds at different prevalence
Figure 5: Variance of Evaluation Metrics by Number of Thresholds.
...and 9 more figures

Area under the ROC Curve has the Most Consistent Evaluation for Binary Classification

TL;DR

Abstract

Area under the ROC Curve has the Most Consistent Evaluation for Binary Classification

Authors

TL;DR

Abstract

Table of Contents

Figures (14)