Dynamical model parameters from ultrasound tongue kinematics
Sam Kirkham, Patrycja Strycharczuk
TL;DR
This study tests whether parameters of a linear harmonic oscillator with a moving target $T$, described by $\ddot{x} + b\dot{x} + k(x-T) = 0$, can be estimated from ultrasound tongue kinematics and compared to EMA. Using simultaneous EMA and ultrasound during British vowel production, the authors extract tongue dorsum and jaw trajectories, preprocess with landmark tracking, and fit the oscillator via constrained least squares, comparing parameter estimates across modalities. They report strong fits ($R^2 \ge 0.9$) for both modalities and find that while $k$ and $b$ show high variability with no consistent modality bias, the target parameter $T$ exhibits systematic differences attributable to ultrasound knot tracking and measurement dimensions. The results support using ultrasound to evaluate dynamical articulatory models, with the added benefit of tracking posterior tongue regions and jaw dynamics through short tendon signals, enabling broader and less invasive data collection for articulatory theory and applications.
Abstract
The control of speech can be modelled as a dynamical system in which articulators are driven toward target positions. These models are typically evaluated using fleshpoint data, such as electromagnetic articulography (EMA), but recent methodological advances make ultrasound imaging a promising alternative. We evaluate whether the parameters of a linear harmonic oscillator can be reliably estimated from ultrasound tongue kinematics and compare these with parameters estimated from simultaneously-recorded EMA data. We find that ultrasound and EMA yield comparable dynamical parameters, while mandibular short tendon tracking also adequately captures jaw motion. This supports using ultrasound kinematics to evaluate dynamical articulatory models.
