Table of Contents
Fetching ...

On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy

Aline Mangold, Juliane Zietz, Susanne Weinhold, Sebastian Pannasch

TL;DR

This paper tackles the lack of human-centered evaluation in explainable AI by conducting a composite literature review of 65 user studies across domains. It develops a three-tier taxonomy separating the core XAI system, its explanations, and the user, and categorizes evaluation metrics into affection, cognition, usability, and interpretability, with additional metrics for explanations. The authors extend existing design frameworks to include AI novices and data experts, outlining design goals such as responsible use, acceptance, usability, human–AI collaboration, and system and task performance. They offer guidelines for validation, methodology reporting, holistic evaluation, explanation assessment, and considering behavioral intentions to improve the rigor and relevance of XAI studies. Collectively, the work provides a comprehensive, human-centered blueprint to guide future XAI development, evaluation, and standardization, boosting the practical impact of explainable systems.

Abstract

As AI becomes more common in everyday living, there is an increasing demand for intelligent systems that are both performant and understandable. Explainable AI (XAI) systems aim to provide comprehensible explanations of decisions and predictions. At present, however, evaluation processes are rather technical and not sufficiently focused on the needs of human users. Consequently, evaluation studies involving human users can serve as a valuable guide for conducting user studies. This paper presents a comprehensive review of 65 user studies evaluating XAI systems across different domains and application contexts. As a guideline for XAI developers, we provide a holistic overview of the properties of XAI systems and evaluation metrics focused on human users (human-centered). We propose objectives for the human-centered design (design goals) of XAI systems. To incorporate users' specific characteristics, design goals are adapted to users with different levels of AI expertise (AI novices and data experts). In this regard, we provide an extension to existing XAI evaluation and design frameworks. The first part of our results includes the analysis of XAI system characteristics. An important finding is the distinction between the core system and the XAI explanation, which together form the whole system. Further results include the distinction of evaluation metrics into affection towards the system, cognition, usability, interpretability, and explanation metrics. Furthermore, the users, along with their specific characteristics and behavior, can be assessed. For AI novices, the relevant extended design goals include responsible use, acceptance, and usability. For data experts, the focus is performance-oriented and includes human-AI collaboration and system and user task performance.

On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy

TL;DR

This paper tackles the lack of human-centered evaluation in explainable AI by conducting a composite literature review of 65 user studies across domains. It develops a three-tier taxonomy separating the core XAI system, its explanations, and the user, and categorizes evaluation metrics into affection, cognition, usability, and interpretability, with additional metrics for explanations. The authors extend existing design frameworks to include AI novices and data experts, outlining design goals such as responsible use, acceptance, usability, human–AI collaboration, and system and task performance. They offer guidelines for validation, methodology reporting, holistic evaluation, explanation assessment, and considering behavioral intentions to improve the rigor and relevance of XAI studies. Collectively, the work provides a comprehensive, human-centered blueprint to guide future XAI development, evaluation, and standardization, boosting the practical impact of explainable systems.

Abstract

As AI becomes more common in everyday living, there is an increasing demand for intelligent systems that are both performant and understandable. Explainable AI (XAI) systems aim to provide comprehensible explanations of decisions and predictions. At present, however, evaluation processes are rather technical and not sufficiently focused on the needs of human users. Consequently, evaluation studies involving human users can serve as a valuable guide for conducting user studies. This paper presents a comprehensive review of 65 user studies evaluating XAI systems across different domains and application contexts. As a guideline for XAI developers, we provide a holistic overview of the properties of XAI systems and evaluation metrics focused on human users (human-centered). We propose objectives for the human-centered design (design goals) of XAI systems. To incorporate users' specific characteristics, design goals are adapted to users with different levels of AI expertise (AI novices and data experts). In this regard, we provide an extension to existing XAI evaluation and design frameworks. The first part of our results includes the analysis of XAI system characteristics. An important finding is the distinction between the core system and the XAI explanation, which together form the whole system. Further results include the distinction of evaluation metrics into affection towards the system, cognition, usability, interpretability, and explanation metrics. Furthermore, the users, along with their specific characteristics and behavior, can be assessed. For AI novices, the relevant extended design goals include responsible use, acceptance, and usability. For data experts, the focus is performance-oriented and includes human-AI collaboration and system and user task performance.
Paper Structure (37 sections, 7 figures, 2 tables)

This paper contains 37 sections, 7 figures, 2 tables.

Figures (7)

  • Figure 1: User Groups of XAI adapted from Mohseni mohseniMultidisciplinarySurveyFramework2021.
  • Figure 2: Composite Literature Review Method. Reproduced from Brendel brendelWhatLiteratureReview2020, licensed under https://creativecommons.org/licenses/by-nc-nd/4.0/.
  • Figure 3: Development of the Concept Matrix
  • Figure 4: Number of Reviewed Papers by Year.
  • Figure 5: Application Domains among Reviewed Papers.
  • ...and 2 more figures