Table of Contents
Fetching ...

Speculative Model Risk in Healthcare AI: Using Storytelling to Surface Unintended Harms

Xingmeng Zhao, Dan Schumacher, Veronica Rammouz, Anthony Rios

TL;DR

The paper tackles the challenge of surfacing unintended harms in healthcare AI by combining automated generation of context-rich user stories with multi-agent red-team discussions. It introduces a storytelling-driven framework that uses a language-based world model and role-playing prompts to foster human-centered ethical foresight before deployment. Empirical evaluation shows that story-driven methods yield broader, more diverse harm and benefit reasoning, with ablations demonstrating the importance of environment progression and participant perspectives. The work highlights the value of narrative scaffolds over automated risk prediction alone for improving ethical reasoning in AI healthcare design and suggests directions for integrating storytelling into existing risk-assessment workflows.

Abstract

Artificial intelligence (AI) is rapidly transforming healthcare, enabling fast development of tools like stress monitors, wellness trackers, and mental health chatbots. However, rapid and low-barrier development can introduce risks of bias, privacy violations, and unequal access, especially when systems ignore real-world contexts and diverse user needs. Many recent methods use AI to detect risks automatically, but this can reduce human engagement in understanding how harms arise and who they affect. We present a human-centered framework that generates user stories and supports multi-agent discussions to help people think creatively about potential benefits and harms before deployment. In a user study, participants who read stories recognized a broader range of harms, distributing their responses more evenly across all 13 harm types. In contrast, those who did not read stories focused primarily on privacy and well-being (58.3%). Our findings show that storytelling helped participants speculate about a broader range of harms and benefits and think more creatively about AI's impact on users.

Speculative Model Risk in Healthcare AI: Using Storytelling to Surface Unintended Harms

TL;DR

The paper tackles the challenge of surfacing unintended harms in healthcare AI by combining automated generation of context-rich user stories with multi-agent red-team discussions. It introduces a storytelling-driven framework that uses a language-based world model and role-playing prompts to foster human-centered ethical foresight before deployment. Empirical evaluation shows that story-driven methods yield broader, more diverse harm and benefit reasoning, with ablations demonstrating the importance of environment progression and participant perspectives. The work highlights the value of narrative scaffolds over automated risk prediction alone for improving ethical reasoning in AI healthcare design and suggests directions for integrating storytelling into existing risk-assessment workflows.

Abstract

Artificial intelligence (AI) is rapidly transforming healthcare, enabling fast development of tools like stress monitors, wellness trackers, and mental health chatbots. However, rapid and low-barrier development can introduce risks of bias, privacy violations, and unequal access, especially when systems ignore real-world contexts and diverse user needs. Many recent methods use AI to detect risks automatically, but this can reduce human engagement in understanding how harms arise and who they affect. We present a human-centered framework that generates user stories and supports multi-agent discussions to help people think creatively about potential benefits and harms before deployment. In a user study, participants who read stories recognized a broader range of harms, distributing their responses more evenly across all 13 harm types. In contrast, those who did not read stories focused primarily on privacy and well-being (58.3%). Our findings show that storytelling helped participants speculate about a broader range of harms and benefits and think more creatively about AI's impact on users.
Paper Structure (16 sections, 1 equation, 17 figures, 10 tables)

This paper contains 16 sections, 1 equation, 17 figures, 10 tables.

Figures (17)

  • Figure 1: Illustration of using speculative stories to help people imagine potential harms and benefits of healthcare AI and foster more creative and ethical thinking.
  • Figure 2: Overview of the Storytelling Framework. We first generate use case scenarios from AI concepts sourced from PubMed, Wired, and industry app descriptions. Next, we simulates role-playing and environment trajectories for each scenario, producing detailed simulation logs. Finally, we rephrase these logs into short stories that illustrate both potential benefits and harms of the AI system.
  • Figure 3: Results of human preference evaluation. Our Storytelling method achieves strong preference wins against the baseline, with 88% preference using Llama3 and 76% using Gemma3.
  • Figure 4: A qualitative example showing how our storytelling method makes the AI’s decision process and its consequences easy to follow. Unlike a simple narrative description, the story explicitly surfaces what changed, why it changed, and how stakeholders were affected.
  • Figure 5: A comparison of simulation logs under different ablations (w/o Environment Trajectories and w/o Role-Playing) to show the contribution of each component.
  • ...and 12 more figures