Table of Contents
Fetching ...

Beyond the Voice: Inertial Sensing of Mouth Motion for High Security Speech Verification

Ynes Ineza, Muhammad A. Ullah, Abdul Serwadda, Aurore Munyaneza

TL;DR

This work addresses the growing vulnerability of voice authentication to deepfake attacks by introducing inertial mouth-movement as a second authentication factor. A chin-mounted inertial sensor array around the mouth captures vertical jaw displacement and lateral facial dynamics, enabling a multimodal system that fuses motion signals with audio for real-time verification. In experiments with 43 participants across seated, walking, and stair-climbing conditions and including native and non-native English speakers, a two-layer LSTM using the chin signal achieves median equal error rates (EER) below $0.01$, while cheek sensors offer marginal gains for margin-based models. A rigorous video-driven attack evaluation shows that publicly available video cannot reliably spoof the inertial mouth-movement biometrics (FAR = $0\%$) under realistic video quality, underscoring the practicality of this approach as a lightweight, deployable defense against voice forgery.

Abstract

Voice interfaces are increasingly used in high stakes domains such as mobile banking, smart home security, and hands free healthcare. Meanwhile, modern generative models have made high quality voice forgeries inexpensive and easy to create, eroding confidence in voice authentication alone. To strengthen protection against such attacks, we present a second authentication factor that combines acoustic evidence with the unique motion patterns of a speaker's lower face. By placing lightweight inertial sensors around the mouth to capture mouth opening and evolving lower facial geometry, our system records a distinct motion signature with strong discriminative power across individuals. We built a prototype and recruited 43 participants to evaluate the system under four conditions seated, walking on level ground, walking on stairs, and speaking with different language backgrounds (native vs. non native English). Across all scenarios, our approach consistently achieved a median equal error rate (EER) of 0.01 or lower, indicating that mouth movement data remain robust under variations in gait, posture, and spoken language. We discuss specific use cases where this second line of defense could provide tangible security benefits to voice authentication systems.

Beyond the Voice: Inertial Sensing of Mouth Motion for High Security Speech Verification

TL;DR

This work addresses the growing vulnerability of voice authentication to deepfake attacks by introducing inertial mouth-movement as a second authentication factor. A chin-mounted inertial sensor array around the mouth captures vertical jaw displacement and lateral facial dynamics, enabling a multimodal system that fuses motion signals with audio for real-time verification. In experiments with 43 participants across seated, walking, and stair-climbing conditions and including native and non-native English speakers, a two-layer LSTM using the chin signal achieves median equal error rates (EER) below , while cheek sensors offer marginal gains for margin-based models. A rigorous video-driven attack evaluation shows that publicly available video cannot reliably spoof the inertial mouth-movement biometrics (FAR = ) under realistic video quality, underscoring the practicality of this approach as a lightweight, deployable defense against voice forgery.

Abstract

Voice interfaces are increasingly used in high stakes domains such as mobile banking, smart home security, and hands free healthcare. Meanwhile, modern generative models have made high quality voice forgeries inexpensive and easy to create, eroding confidence in voice authentication alone. To strengthen protection against such attacks, we present a second authentication factor that combines acoustic evidence with the unique motion patterns of a speaker's lower face. By placing lightweight inertial sensors around the mouth to capture mouth opening and evolving lower facial geometry, our system records a distinct motion signature with strong discriminative power across individuals. We built a prototype and recruited 43 participants to evaluate the system under four conditions seated, walking on level ground, walking on stairs, and speaking with different language backgrounds (native vs. non native English). Across all scenarios, our approach consistently achieved a median equal error rate (EER) of 0.01 or lower, indicating that mouth movement data remain robust under variations in gait, posture, and spoken language. We discuss specific use cases where this second line of defense could provide tangible security benefits to voice authentication systems.
Paper Structure (40 sections, 2 equations, 8 figures, 4 tables)

This paper contains 40 sections, 2 equations, 8 figures, 4 tables.

Figures (8)

  • Figure 1: Operational flow of continuous authentication using mouth-motion signals
  • Figure 2: User wearing our authentication prototype
  • Figure 3: Motion sensor used in our experiment
  • Figure 4: Mean acceleration in the X, Y and Z directions for 3 users speaking continuously for 30 seconds. Each of the users spoke different content from the other users.
  • Figure 5: Mean acceleration in the X, Y and Z directions for 3 users who read from a common script. All users spoke exactly the same content.
  • ...and 3 more figures