Table of Contents
Fetching ...

Dynamical model parameters from ultrasound tongue kinematics

Sam Kirkham, Patrycja Strycharczuk

TL;DR

This study tests whether parameters of a linear harmonic oscillator with a moving target $T$, described by $\ddot{x} + b\dot{x} + k(x-T) = 0$, can be estimated from ultrasound tongue kinematics and compared to EMA. Using simultaneous EMA and ultrasound during British vowel production, the authors extract tongue dorsum and jaw trajectories, preprocess with landmark tracking, and fit the oscillator via constrained least squares, comparing parameter estimates across modalities. They report strong fits ($R^2 \ge 0.9$) for both modalities and find that while $k$ and $b$ show high variability with no consistent modality bias, the target parameter $T$ exhibits systematic differences attributable to ultrasound knot tracking and measurement dimensions. The results support using ultrasound to evaluate dynamical articulatory models, with the added benefit of tracking posterior tongue regions and jaw dynamics through short tendon signals, enabling broader and less invasive data collection for articulatory theory and applications.

Abstract

The control of speech can be modelled as a dynamical system in which articulators are driven toward target positions. These models are typically evaluated using fleshpoint data, such as electromagnetic articulography (EMA), but recent methodological advances make ultrasound imaging a promising alternative. We evaluate whether the parameters of a linear harmonic oscillator can be reliably estimated from ultrasound tongue kinematics and compare these with parameters estimated from simultaneously-recorded EMA data. We find that ultrasound and EMA yield comparable dynamical parameters, while mandibular short tendon tracking also adequately captures jaw motion. This supports using ultrasound kinematics to evaluate dynamical articulatory models.

Dynamical model parameters from ultrasound tongue kinematics

TL;DR

This study tests whether parameters of a linear harmonic oscillator with a moving target , described by , can be estimated from ultrasound tongue kinematics and compared to EMA. Using simultaneous EMA and ultrasound during British vowel production, the authors extract tongue dorsum and jaw trajectories, preprocess with landmark tracking, and fit the oscillator via constrained least squares, comparing parameter estimates across modalities. They report strong fits () for both modalities and find that while and show high variability with no consistent modality bias, the target parameter exhibits systematic differences attributable to ultrasound knot tracking and measurement dimensions. The results support using ultrasound to evaluate dynamical articulatory models, with the added benefit of tracking posterior tongue regions and jaw dynamics through short tendon signals, enabling broader and less invasive data collection for articulatory theory and applications.

Abstract

The control of speech can be modelled as a dynamical system in which articulators are driven toward target positions. These models are typically evaluated using fleshpoint data, such as electromagnetic articulography (EMA), but recent methodological advances make ultrasound imaging a promising alternative. We evaluate whether the parameters of a linear harmonic oscillator can be reliably estimated from ultrasound tongue kinematics and compare these with parameters estimated from simultaneously-recorded EMA data. We find that ultrasound and EMA yield comparable dynamical parameters, while mandibular short tendon tracking also adequately captures jaw motion. This supports using ultrasound kinematics to evaluate dynamical articulatory models.
Paper Structure (12 sections, 3 equations, 4 figures, 2 tables)

This paper contains 12 sections, 3 equations, 4 figures, 2 tables.

Figures (4)

  • Figure 1: Left: Location of DLC knots estimated for each ultrasound frame (knot 1 is tongue root, knot 11 is tongue tip, H is hyoid, M is mandible, ST is short tendon). Right: Raw and smoothed data for TD horizontal position from EMA and ultrasound in the word bar.
  • Figure 2: A random sample of three example velocity fits between EMA and Ultrasound for TDx. Tokens were selected using a fixed random seed and each word represents the same underlying token produced by a speaker. All fits are $R^2 > 0.92$.
  • Figure 3: By-word effects showing how ultrasound-estimated $T$ (target) values differ from EMA-estimated values. Word labels show estimated means; blue lines show 95% credible intervals. A value of zero indicates that ultrasound parameters do not differ from EMA parameters.
  • Figure 4: By-word effects showing how ultrasound-estimated $k$ (stiffness) and $b$ (damping) values differ from EMA-estimated values. Word labels show estimated means; blue lines show 95% credible intervals. A value of zero indicates that ultrasound parameters do not differ from EMA parameters. Note that the word bar has been removed from the JAW plots due to excessively large confidence intervals (crossing zero in both dimensions) that distort the axis ranges.