3D markerless tracking of speech movements with submillimeter accuracy.
Journal:
IEEE transactions on neural systems and rehabilitation engineering : a publication of the IEEE Engineering in Medicine and Biology Society
Published Date:
Apr 30, 2026
Abstract
Speech movements, generated through precise spatial and temporal control of articulators, are inherently complex. Measuring these movements is challenging and often requires the use of multiple physical sensors positioned around the mouth and face to acquire precise movement measurements. Facial sensor placement can be difficult for certain populations to tolerate, particularly young children. Recent progress in machine learning-based markerless facial landmark tracking technology has demonstrated potential to provide lip tracking without the need for physical sensors, but whether such technology can provide submillimeter precision and accuracy in 3D remains unclear. Here, we developed a novel approach that integrates a facial landmark detector and CoTracker, a transformer-based neural network model that jointly tracks dense points across a video sequence. We further examined and validated this approach by assessing its tracking precision and accuracy. The findings revealed that our approach was more precise (≈ 0.15 mm in standard deviation) than a facial landmark detector alone (> 0.3 mm). In addition, its 3D tracking performance was comparable to electromagnetic articulography (≈ 0.3 mm RMSE against simultaneously recorded articulograph data). Importantly, the approach performed similarly well across adults and young children (i.e., 3- and 4-year-olds). Our novel framework leverages open-source pre-trained models, promoting accessibility and open science while using commercial-grade compute resources. It also serves as a proof of concept for improving the performance of a broad range of commonly used markerless tracking applications in neuroscience.
Authors
Keywords
No keywords available for this article.