Alignment of Artificial Neural Networks with Human Eye Movement Patterns in Facial Expression Recognition
Journal:
bioRxiv
Published Date:
Jan 1, 2025
Abstract
Artificial neural networks (ANNs) have become powerful tools for modeling human perception and hehavior, yet it remains unclear if they resemble the sequential eye movement strategies humans employ during visual processing. Human studies demonstrate that the first two fixations play distinct roles in face processing: the initial fixation (Fix I) captures global configurational information, while the second (Fix II) targets local diagnostic details, with the two working collaboratively to support recognition. In this study, we tested if such properties emerge in ANNs trained for facial expression recognition. We fine-tuned a convolutional neural network (VGG16) and a Vision Transformer (ViT) on multiple facial expression datasets and evaluated their performance on the same images used in a human eye-tracking experiment. Human participants (N = 28) achieved a mean accuracy of 84.2%, VGG reached 82.9%, and ViT outperformed both at 90.0%. Perturbation analyses showed that both models aligned with the independent functions of Fix I and Fix II, but only ViT displayed a synergistic effect when information from both fixations was combined. Layer-wise analyses revealed stronger correspondence with Fix II than Fix I in both models, while ViT uniquely captured the statistical dependencies of fixation transitions. These findings suggest that VGG primarily reflects feedforward, detail-based processing, whereas ViT supports feedback-like integration across sequential fixations. Our results indicate that current ANNs can approximate human eye movement patterns in face processing, and highlight the potential of transformer-based architectures to model sequential strategies in human visual exploration.