Table of Contents
Fetching ...

Assessing the Real-World Utility of Explainable AI for Arousal Diagnostics: An Application-Grounded User Study

Stefan Kraft, Andreas Theissler, Vera Wienhausen-Wilke, Gjergji Kasneci, Hendrik Lensch

TL;DR

This study probes the real-world value of explainable AI for arousal diagnostics in polysomnography by conducting an application-grounded user study with eight professional sleep scorers. It compares manual scoring, black-box AI, and transparent white-box AI across start and quality-control timing, using both consensus and CPS-ground-truth benchmarks. The results show that WB explanations plus a QC workflow substantially improve count-based and, to a lesser extent, event-level performance, while WB also enhances trust and perceived usefulness, albeit with higher time demands. Importantly, outcomes depend on the ground-truth standard used, highlighting the need for carefully chosen training standards and governance when deploying clinical decision-support systems. Overall, strategically timed transparent AI serves as an effective co-scorer, balancing accuracy, efficiency, and user acceptance for potential clinical adoption.

Abstract

Artificial intelligence (AI) systems increasingly match or surpass human experts in biomedical signal interpretation. However, their effective integration into clinical practice requires more than high predictive accuracy. Clinicians must discern \textit{when} and \textit{why} to trust algorithmic recommendations. This work presents an application-grounded user study with eight professional sleep medicine practitioners, who score nocturnal arousal events in polysomnographic data under three conditions: (i) manual scoring, (ii) black-box (BB) AI assistance, and (iii) transparent white-box (WB) AI assistance. Assistance is provided either from the \textit{start} of scoring or as a post-hoc quality-control (\textit{QC}) review. We systematically evaluate how the type and timing of assistance influence event-level and clinically most relevant count-based performance, time requirements, and user experience. When evaluated against the clinical standard used to train the AI, both AI and human-AI teams significantly outperform unaided experts, with collaboration also reducing inter-rater variability. Notably, transparent AI assistance applied as a targeted QC step yields median event-level performance improvements of approximately 30\% over black-box assistance, and QC timing further enhances count-based outcomes. While WB and QC approaches increase the time required for scoring, start-time assistance is faster and preferred by most participants. Participants overwhelmingly favor transparency, with seven out of eight expressing willingness to adopt the system with minor or no modifications. In summary, strategically timed transparent AI assistance effectively balances accuracy and clinical efficiency, providing a promising pathway toward trustworthy AI integration and user acceptance in clinical workflows.

Assessing the Real-World Utility of Explainable AI for Arousal Diagnostics: An Application-Grounded User Study

TL;DR

This study probes the real-world value of explainable AI for arousal diagnostics in polysomnography by conducting an application-grounded user study with eight professional sleep scorers. It compares manual scoring, black-box AI, and transparent white-box AI across start and quality-control timing, using both consensus and CPS-ground-truth benchmarks. The results show that WB explanations plus a QC workflow substantially improve count-based and, to a lesser extent, event-level performance, while WB also enhances trust and perceived usefulness, albeit with higher time demands. Importantly, outcomes depend on the ground-truth standard used, highlighting the need for carefully chosen training standards and governance when deploying clinical decision-support systems. Overall, strategically timed transparent AI serves as an effective co-scorer, balancing accuracy, efficiency, and user acceptance for potential clinical adoption.

Abstract

Artificial intelligence (AI) systems increasingly match or surpass human experts in biomedical signal interpretation. However, their effective integration into clinical practice requires more than high predictive accuracy. Clinicians must discern \textit{when} and \textit{why} to trust algorithmic recommendations. This work presents an application-grounded user study with eight professional sleep medicine practitioners, who score nocturnal arousal events in polysomnographic data under three conditions: (i) manual scoring, (ii) black-box (BB) AI assistance, and (iii) transparent white-box (WB) AI assistance. Assistance is provided either from the \textit{start} of scoring or as a post-hoc quality-control (\textit{QC}) review. We systematically evaluate how the type and timing of assistance influence event-level and clinically most relevant count-based performance, time requirements, and user experience. When evaluated against the clinical standard used to train the AI, both AI and human-AI teams significantly outperform unaided experts, with collaboration also reducing inter-rater variability. Notably, transparent AI assistance applied as a targeted QC step yields median event-level performance improvements of approximately 30\% over black-box assistance, and QC timing further enhances count-based outcomes. While WB and QC approaches increase the time required for scoring, start-time assistance is faster and preferred by most participants. Participants overwhelmingly favor transparency, with seven out of eight expressing willingness to adopt the system with minor or no modifications. In summary, strategically timed transparent AI assistance effectively balances accuracy and clinical efficiency, providing a promising pathway toward trustworthy AI integration and user acceptance in clinical workflows.
Paper Structure (134 sections, 9 equations, 25 figures, 18 tables)

This paper contains 134 sections, 9 equations, 25 figures, 18 tables.

Figures (25)

  • Figure 1: Main interface of the Decision Support System for the arousal scoring task. The interface is organized into several key areas, each marked with blue circles: 1) Control bar providing access to basic information, and navigation tools; 2) Overview graph displaying the hypnogram (sleep stages) for the entire recording, with the currently selected interval highlighted by an orange vertical bar; 3) Timeline summarizing AI-annotated arousal event positions across the recording; 4) Visualization of selected polysomnography data channels for the current interval, including overlays for AI-suggested arousal regions; 5) Event information panel presenting details of the currently selected arousal event, including options to accept or reject the event; Additional transparency elements, highlighted with orange ovals, include: 4a) Confidence score channel, visualizing the AI model's confidence for arousal onset at each time point; 4b) Shaded regions on each channel, indicating intervals where an arousal onset is likely, with the most probable onset marked by a vertical dashed line (corresponding to the maximum confidence score); 5a) Bar indicator showing the maximum confidence score for the selected event, visualized with a green gradient.
  • Figure 2: Local explanations of the AI model's arousal prediction at three levels of detail for the arousal depicted in Figure \ref{['fig:dss-main-interface']}. Each panel visualizes the ten most relevant channels for a predicted arousal onset, with the x-axis representing time relative to the suggested arousal onset (vertical line at $t=0$). Blue dots indicate data points with relevance scores above the current threshold, and the color bar to the right encodes the magnitude of feature relevance (dark green: high relevance, light green: lower relevance). Gray dashed lines mark the interval boundaries where an arousal is most likely to start. The top panel shows the medium detail level, while the bottom left and bottom right panels display the low and high detail levels, respectively, corresponding to different thresholds for feature attribution. Legends clarify the graphical elements.
  • Figure 3: Global explanation for the AI model's decision making process. The bar chart displays the global relevance percentage of each channel for the AI model's arousal prediction, aggregated across the dataset. Channels are ranked in descending order of their contribution, with RIP Abdomen, Pulse, and RIP Thorax showing the highest relevance. This visualization helps users understand which physiological signals most strongly influence the model's decisions at a global level.
  • Figure 4: Detailed demographic profile of study participants, illustrating basic demographic information (orange), professional experience in sleep diagnostics (blue), and both AI experience and motivation for participation (light blue).
  • Figure 5: Overview of the Study Phases. The phases are annotated in orange. Dashed lines indicate that the next phase is only started after all steps of the previous phase are completed. Abbreviations: QC denotes Quality Control, WB refers to White Box, and BB means Black Box.
  • ...and 20 more figures