Table of Contents
Fetching ...

What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics

Lennart Wachowiak, Andrew Coles, Gerard Canal, Oya Celiktutan

TL;DR

This work presents a large-scale, empirically grounded dataset of 1,893 user questions about household robots, collected from 100 participants across videos and text scenarios to illuminate what information a robot should be able to explain. The authors develop a two-level hierarchical taxonomy of 12 main categories and 70 subcategories (per their coding scheme) and analyze how question importance varies by category, as well as by user attitudes toward robots and robotics experience using linear mixed-effects models and regression. Key findings show that questions about potential issues and how robots know or decide things are prioritized by users, while why-questions and mental-state inquiries are rated lower, with attitudes and experience shaping what users ask. These insights guide data logging, benchmarking QA/NL explainability modules, and designing explanations aligned with user expectations, while acknowledging limitations such as ecological validity and cross-cultural generalizability.

Abstract

With the growing use of large language models and conversational interfaces in human-robot interaction, robots' ability to answer user questions is more important than ever. We therefore introduce a dataset of 1,893 user questions for household robots, collected from 100 participants and organized into 12 categories and 70 subcategories. Most work in explainable robotics focuses on why-questions. In contrast, our dataset provides a wide variety of questions, from questions about simple execution details to questions about how the robot would act in hypothetical scenarios -- thus giving roboticists valuable insights into what questions their robot needs to be able to answer. To collect the dataset, we created 15 video stimuli and 7 text stimuli, depicting robots performing varied household tasks. We then asked participants on Prolific what questions they would want to ask the robot in each portrayed situation. In the final dataset, the most frequent categories are questions about task execution details (22.5%), the robot's capabilities (12.7%), and performance assessments (11.3%). Although questions about how robots would handle potentially difficult scenarios and ensure correct behavior are less frequent, users rank them as the most important for robots to be able to answer. Moreover, we find that users who identify as novices in robotics ask different questions than more experienced users. Novices are more likely to inquire about simple facts, such as what the robot did or the current state of the environment. As robots enter environments shared with humans and language becomes central to giving instructions and interaction, this dataset provides a valuable foundation for (i) identifying the information robots need to log and expose to conversational interfaces, (ii) benchmarking question-answering modules, and (iii) designing explanation strategies that align with user expectations.

What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics

TL;DR

This work presents a large-scale, empirically grounded dataset of 1,893 user questions about household robots, collected from 100 participants across videos and text scenarios to illuminate what information a robot should be able to explain. The authors develop a two-level hierarchical taxonomy of 12 main categories and 70 subcategories (per their coding scheme) and analyze how question importance varies by category, as well as by user attitudes toward robots and robotics experience using linear mixed-effects models and regression. Key findings show that questions about potential issues and how robots know or decide things are prioritized by users, while why-questions and mental-state inquiries are rated lower, with attitudes and experience shaping what users ask. These insights guide data logging, benchmarking QA/NL explainability modules, and designing explanations aligned with user expectations, while acknowledging limitations such as ecological validity and cross-cultural generalizability.

Abstract

With the growing use of large language models and conversational interfaces in human-robot interaction, robots' ability to answer user questions is more important than ever. We therefore introduce a dataset of 1,893 user questions for household robots, collected from 100 participants and organized into 12 categories and 70 subcategories. Most work in explainable robotics focuses on why-questions. In contrast, our dataset provides a wide variety of questions, from questions about simple execution details to questions about how the robot would act in hypothetical scenarios -- thus giving roboticists valuable insights into what questions their robot needs to be able to answer. To collect the dataset, we created 15 video stimuli and 7 text stimuli, depicting robots performing varied household tasks. We then asked participants on Prolific what questions they would want to ask the robot in each portrayed situation. In the final dataset, the most frequent categories are questions about task execution details (22.5%), the robot's capabilities (12.7%), and performance assessments (11.3%). Although questions about how robots would handle potentially difficult scenarios and ensure correct behavior are less frequent, users rank them as the most important for robots to be able to answer. Moreover, we find that users who identify as novices in robotics ask different questions than more experienced users. Novices are more likely to inquire about simple facts, such as what the robot did or the current state of the environment. As robots enter environments shared with humans and language becomes central to giving instructions and interaction, this dataset provides a valuable foundation for (i) identifying the information robots need to log and expose to conversational interfaces, (ii) benchmarking question-answering modules, and (iii) designing explanation strategies that align with user expectations.
Paper Structure (44 sections, 2 equations, 5 figures, 4 tables)

This paper contains 44 sections, 2 equations, 5 figures, 4 tables.

Figures (5)

  • Figure 1: Participant Statistics: their experience with robots (Likert scale 1--7), personal/societal-level attitudes towards robots (measured via GAToRS koverola2022general, an average of multiple items using a 1--7 Likert scale), age, gender, ethnicity, and binary statistics on whether they programmed a robot before, interacted with one, or studied a computer science or engineering.
  • Figure 2: Hierarchical categorization of user questions for the robot (see Tables \ref{['tab:codes']}, \ref{['tab:codes2']}, and \ref{['tab:codes3']} for category definitions).
  • Figure 3: Estimated marginal means (95% CIs) of the importance scores per question category --- obtained via a linear mixed-effects model. Number of samples per category (n) shown on the right, and statistical significance indicated by: *** = $p<.001$, ** = $p<.01$, * = $p<.05$
  • Figure 4: Users' rating of how important they think it is that a robot can answer their proposed questions is related to their general attitudes towards robots as measured by the GAToRS questionnaire.
  • Figure 5: Difference in question type distribution based on robot experience. Robot experience is based on self-report on a Likert scale 1--7. There are 76 participants with a score between 1 and 3, and 24 participants with a score between 4 and 7.