Table of Contents
Fetching ...

Finding Holes: Pathologist Level Performance Using AI for Cribriform Morphology Detection in Prostate Cancer

Kelvin Szolnoky, Anders Blilie, Nita Mulliqi, Toyonori Tsuzuki, Hemamali Samaratunga, Matteo Titus, Xiaoyi Ji, Sol Erika Boman, Einar Gudlaugsson, Svein Reidar Kjosavik, José Asenjo, Marcello Gambacorta, Paolo Libretti, Marcin Braun, Radisław Kordek, Roman Łowicki, Brett Delahunt, Kenneth A. Iczkowski, Theo van der Kwast, Geert J. L. H. van Leenders, Katia R. M. Leite, Chin-Chen Pan, Emiel Adrianus Maria Janssen, Martin Eklund, Lars Egevad, Kimmo Kartasalo

TL;DR

The study addresses the underdiagnosis and interobserver variability of cribriform morphology in prostate cancer by developing a MIL-based AI system using an EfficientNetV2-S backbone. It leverages multi-cohort training and validates internally and externally across independent scanners, achieving an internal AUC of $0.97$ and external AUC of $0.90$, with corresponding Cohen's kappas of $0.81$ and $0.55$, respectively. The model demonstrates pathologist-level performance and even highest average agreement compared with nine expert pathologists, though external calibration shifts lead to some overdiagnosis at the standard threshold. This approach could standardise cribriform reporting, aid risk stratification, and prioritize diagnostically challenging regions for expert review, representing a meaningful step toward integrating AI assistance into prostate cancer pathology workflows.

Abstract

Background: Cribriform morphology in prostate cancer is a histological feature that indicates poor prognosis and contraindicates active surveillance. However, it remains underreported and subject to significant interobserver variability amongst pathologists. We aimed to develop and validate an AI-based system to improve cribriform pattern detection. Methods: We created a deep learning model using an EfficientNetV2-S encoder with multiple instance learning for end-to-end whole-slide classification. The model was trained on 640 digitised prostate core needle biopsies from 430 patients, collected across three cohorts. It was validated internally (261 slides from 171 patients) and externally (266 slides, 104 patients from three independent cohorts). Internal validation cohorts included laboratories or scanners from the development set, while external cohorts used completely independent instruments and laboratories. Annotations were provided by three expert uropathologists with known high concordance. Additionally, we conducted an inter-rater analysis and compared the model's performance against nine expert uropathologists on 88 slides from the internal validation cohort. Results: The model showed strong internal validation performance (AUC: 0.97, 95% CI: 0.95-0.99; Cohen's kappa: 0.81, 95% CI: 0.72-0.89) and robust external validation (AUC: 0.90, 95% CI: 0.86-0.93; Cohen's kappa: 0.55, 95% CI: 0.45-0.64). In our inter-rater analysis, the model achieved the highest average agreement (Cohen's kappa: 0.66, 95% CI: 0.57-0.74), outperforming all nine pathologists whose Cohen's kappas ranged from 0.35 to 0.62. Conclusion: Our AI model demonstrates pathologist-level performance for cribriform morphology detection in prostate cancer. This approach could enhance diagnostic reliability, standardise reporting, and improve treatment decisions for prostate cancer patients.

Finding Holes: Pathologist Level Performance Using AI for Cribriform Morphology Detection in Prostate Cancer

TL;DR

The study addresses the underdiagnosis and interobserver variability of cribriform morphology in prostate cancer by developing a MIL-based AI system using an EfficientNetV2-S backbone. It leverages multi-cohort training and validates internally and externally across independent scanners, achieving an internal AUC of and external AUC of , with corresponding Cohen's kappas of and , respectively. The model demonstrates pathologist-level performance and even highest average agreement compared with nine expert pathologists, though external calibration shifts lead to some overdiagnosis at the standard threshold. This approach could standardise cribriform reporting, aid risk stratification, and prioritize diagnostically challenging regions for expert review, representing a meaningful step toward integrating AI assistance into prostate cancer pathology workflows.

Abstract

Background: Cribriform morphology in prostate cancer is a histological feature that indicates poor prognosis and contraindicates active surveillance. However, it remains underreported and subject to significant interobserver variability amongst pathologists. We aimed to develop and validate an AI-based system to improve cribriform pattern detection. Methods: We created a deep learning model using an EfficientNetV2-S encoder with multiple instance learning for end-to-end whole-slide classification. The model was trained on 640 digitised prostate core needle biopsies from 430 patients, collected across three cohorts. It was validated internally (261 slides from 171 patients) and externally (266 slides, 104 patients from three independent cohorts). Internal validation cohorts included laboratories or scanners from the development set, while external cohorts used completely independent instruments and laboratories. Annotations were provided by three expert uropathologists with known high concordance. Additionally, we conducted an inter-rater analysis and compared the model's performance against nine expert uropathologists on 88 slides from the internal validation cohort. Results: The model showed strong internal validation performance (AUC: 0.97, 95% CI: 0.95-0.99; Cohen's kappa: 0.81, 95% CI: 0.72-0.89) and robust external validation (AUC: 0.90, 95% CI: 0.86-0.93; Cohen's kappa: 0.55, 95% CI: 0.45-0.64). In our inter-rater analysis, the model achieved the highest average agreement (Cohen's kappa: 0.66, 95% CI: 0.57-0.74), outperforming all nine pathologists whose Cohen's kappas ranged from 0.35 to 0.62. Conclusion: Our AI model demonstrates pathologist-level performance for cribriform morphology detection in prostate cancer. This approach could enhance diagnostic reliability, standardise reporting, and improve treatment decisions for prostate cancer patients.
Paper Structure (21 sections, 7 figures, 10 tables)

This paper contains 21 sections, 7 figures, 10 tables.

Figures (7)

  • Figure 1: Overview of the study design. Phase 1 (Development) used subsets of the STHLM3 and SUH cohorts for model training. Phase 2 (Validation) included internal validation on reserved STHLM3/SUH data and external validation on three independent cohorts (AMU, MUL, SCH). Slides were digitised on scanners from multiple vendors and annotated by three pathologists. Numbers in parentheses indicate scanner serial numbers. Serial numbers for the scanners used at SCH are unavailable, but these scanners are distinct from those used in the other cohorts. No scanners in the external cohorts were present in the training data. Performance evaluation included standard metrics (AUC, Cohen's kappa, sensitivity, specificity), inter-rater analysis comparing our model with nine pathologists, cross-scanner reproducibility assessment, and borderline case analysis. Definition of abbreviations: AUC = Area under the receiver operating characteristic curve.
  • Figure 2: (a) Receiver operating characteristic curves showing model performance on internal and external validation sets. (b) Confusion matrix on predictions for the internal validation set (STHLM3 and SUH). (c) Confusion matrix on predictions for the external validation set (AMU, MUL, and SCH).
  • Figure 3: Pathologist concordance analysis comparing the agreement between our model (robot icon) and nine pathologists (physician icon), showing mean pairwise Cohen's kappa values. For each rater, including our model, the mean pairwise Cohen's kappa was calculated against the other pathologists only (the model was excluded from this average calculation). The whiskers indicate the 95% confidence interval. For exact values see Table \ref{['tab:pathologist-concordance']}.
  • Figure A1: Confusion matrices on predictions for cohorts (a) STHLM3 (b) SUH (c) SCH (d) MUL (e) AMU
  • Figure A2: Receiver operating characteristic curves showing model performance for the different cohorts.
  • ...and 2 more figures