VoiceMorph: How AI Voice Morphing Reveals the Boundaries of Auditory Self-Recognition
Kye Shimizu, Minghan Gao, Ananya Ganesh, Pattie Maes
TL;DR
The paper tackles how far AI-driven voice morphing can push the boundary of recognizing one’s own voice, by interpolating between participants’ voices and demographically matched targets. It combines quantitative self-identification ratings and response times with qualitative interviews in a mixed-methods design across 21 adults. The key findings show a perceptual threshold around $35.2 ext{ extpercent}$ morphing, with older participants tolerating more morphing before losing self-recognition, and that larger embedding distances slow decision-making even when thresholds are unaffected. These results have implications for AI ethics, security, and the design of voice-based interfaces as synthetic voices become more prevalent, highlighting the need to protect vulnerable populations and to consider perceptual boundaries in real-world applications.
Abstract
This study investigated auditory self-recognition boundaries using AI voice morphing technology, examining when individuals cease recognizing their own voice. Through controlled morphing between participants' voices and demographically matched targets at 1% increments using a mixed-methods design, we measured self-identification ratings and response times among 21 participants aged 18-64. Results revealed a critical recognition threshold at 35.2% morphing (95% CI [31.4, 38.1]). Older participants tolerated significantly higher morphing levels before losing self-recognition ($β$ = 0.617, p = 0.048), suggesting age-related vulnerabilities. Greater acoustic embedding distances predicted slower decision-making ($r \approx 0.5-0.53, p < 0.05$), with the longest response times for cloned versions of participants' own voices. Qualitative analysis revealed prosodic-based recognition strategies, universal voice manipulation discomfort, and awareness of applications spanning assistive technology to security risks. These findings establish foundational evidence for individual differences in voice morphing detection, with implications for AI ethics and vulnerable population protection as voice synthesis becomes accessible.
