On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
Aline Mangold, Juliane Zietz, Susanne Weinhold, Sebastian Pannasch
TL;DR
This paper tackles the lack of human-centered evaluation in explainable AI by conducting a composite literature review of 65 user studies across domains. It develops a three-tier taxonomy separating the core XAI system, its explanations, and the user, and categorizes evaluation metrics into affection, cognition, usability, and interpretability, with additional metrics for explanations. The authors extend existing design frameworks to include AI novices and data experts, outlining design goals such as responsible use, acceptance, usability, human–AI collaboration, and system and task performance. They offer guidelines for validation, methodology reporting, holistic evaluation, explanation assessment, and considering behavioral intentions to improve the rigor and relevance of XAI studies. Collectively, the work provides a comprehensive, human-centered blueprint to guide future XAI development, evaluation, and standardization, boosting the practical impact of explainable systems.
Abstract
As AI becomes more common in everyday living, there is an increasing demand for intelligent systems that are both performant and understandable. Explainable AI (XAI) systems aim to provide comprehensible explanations of decisions and predictions. At present, however, evaluation processes are rather technical and not sufficiently focused on the needs of human users. Consequently, evaluation studies involving human users can serve as a valuable guide for conducting user studies. This paper presents a comprehensive review of 65 user studies evaluating XAI systems across different domains and application contexts. As a guideline for XAI developers, we provide a holistic overview of the properties of XAI systems and evaluation metrics focused on human users (human-centered). We propose objectives for the human-centered design (design goals) of XAI systems. To incorporate users' specific characteristics, design goals are adapted to users with different levels of AI expertise (AI novices and data experts). In this regard, we provide an extension to existing XAI evaluation and design frameworks. The first part of our results includes the analysis of XAI system characteristics. An important finding is the distinction between the core system and the XAI explanation, which together form the whole system. Further results include the distinction of evaluation metrics into affection towards the system, cognition, usability, interpretability, and explanation metrics. Furthermore, the users, along with their specific characteristics and behavior, can be assessed. For AI novices, the relevant extended design goals include responsible use, acceptance, and usability. For data experts, the focus is performance-oriented and includes human-AI collaboration and system and user task performance.
