"Over-the-Hood" AI Inclusivity Bugs and How 3 AI Product Teams Found and Fixed Them
Andrew Anderson, Fatima A. Moussaoui, Jimena Noa Guevara, Md Montaser Hamid, Margaret Burnett
TL;DR
The paper defines and investigates 'over-the-hood' AI inclusivity bugs—user-facing issues in AI products that disproportionately affect users with specific problem-solving approaches. Through a field study with three AI product teams (Game, Weather, Farm) using GenderMag, it identifies six AI inclusivity bug types across 83 instances and demonstrates fixes for 47 of them, while introducing GenderMag-for-AI variants to address AI-specific challenges. An external study corroborates that post-GenderMag fixes improve users' mental models and gender equity, highlighting practical benefits for real-world AI deployment. The work shows that while the original GenderMag method helps uncover AI inclusivity bugs, AI-specific adaptations—especially a Pre-Action Fork variant—better capture user doubt about AI outputs and guide more effective fixes. Collectively, the findings offer actionable guidance for AI teams to design more inclusive user-facing AI systems by aligning explanations and interactions with diverse problem-solving styles and trust dynamics.
Abstract
While much research has shown the presence of AI's "under-the-hood" biases (e.g., algorithmic, training data, etc.), what about "over-the-hood" inclusivity biases: barriers in user-facing AI products that disproportionately exclude users with certain problem-solving approaches? Recent research has begun to report the existence of such biases -- but what do they look like, how prevalent are they, and how can developers find and fix them? To find out, we conducted a field study with 3 AI product teams, to investigate what kinds of AI inclusivity bugs exist uniquely in user-facing AI products, and whether/how AI product teams might harness an existing (non-AI-oriented) inclusive design method to find and fix them. The teams' work resulted in identifying 6 types of AI inclusivity bugs arising 83 times, fixes covering 47 of these bug instances, and a new variation of the GenderMag inclusive design method, GenderMag-for-AI, that is especially effective at detecting certain kinds of AI inclusivity bugs.
