Using neural networks to understand static and dynamic cues in facial expression recognition.
Journal:
Vision research
Published Date:
Apr 10, 2026
Abstract
Dynamic facial expression recognition (DFER), which uses dynamic sequences (e.g., videos) of facial expressions rather than single static frames, incorporates temporal information about facial movements and thus greater ecological validity compared to static expression recognition. We studied DFER using a large, naturalistic database (Dynamic Facial Expressions in the Wild, DFEW) with over 11,000 video clips of 7 different facial expressions (Happy, Angry, Neutral, Surprise, Sad, Fear, Disgust). We created four modified versions of the database by varying Input type (Static cues or frames vs. Dynamic cues or optical flows) and temporal Structure (Ordered vs. Shuffled). This study examines and compares the facial recognition performance of deep convolutional neural networks (DCNNs) trained on these four Static and Dynamic datasets. Model performance was evaluated and compared across conditions using univariate and multivariate approaches. We found that DCNNs can recognize facial expressions from either Static or Dynamic cues. However, unlike in action recognition, Static cues consistently outperform Dynamic cues, perhaps due to the smaller magnitude of motion cues in facial expressions. Surprisingly, we found that disrupting temporal Structure did not affect the recognition of most facial expressions. Finally, multivariate analyses of the penultimate layers of the DCNNs suggested significant differences in the underlying representation of facial expressions between Static and Dynamic cues. Taken together, these findings provide new insights into the differential roles of Static and Dynamic cues in facial expression recognition.
Authors
Keywords
No keywords available for this article.