Shared texture-like representations underlie deep neural network alignment with human visual processing.

Journal: Current biology : CB
Published Date:

Abstract

Deep neural networks (DNNs) excel at predicting neural responses across the visual hierarchy,1,2,3,4,5 a success widely interpreted as evidence of shared object recognition computations.6,7 Yet improving DNN object recognition accuracy does not reliably increase neural predictivity,8,9 and even untrained networks predict brain responses above chance.9,10,11 This disconnect suggests that object recognition may not drive DNN-brain alignment. Texture-like statistics are represented in both DNNs and mid-level visual cortex (A.V. Jagadeesh and M. Livingstone, 2024, ICLR, presentation).12,13,1416 In natural images, these statistics are carried by objects and backgrounds, shaping representations and recognition in both systems.16,17,18,19 Does DNN-brain alignment reflect a shared sensitivity to object-related information or texture-like statistics? To dissociate these factors, we recorded electroencephalograms (EEGs) from 57 participants viewing natural scenes, texture-synthesized images preserving local statistics while disrupting global form, and object-only images with backgrounds removed. If alignment reflects texture-like statistics, then it should peak for texture-synthesized images. If it reflects object-related processing, then alignment should be strongest for natural and object-only conditions, which preserve object information. We compared EEG responses with DNN activations via weighted representational similarity analysis.20,21 Texture-synthesized images yielded the strongest DNN-EEG alignment, peaking in early responses (<200 ms) and explaining up to ∼85% of noise-ceiling-normalized explainable variance versus ∼44% for natural and ∼55% for isolated objects. Crucially, object categories were more decodable for natural and object-only images than texture-synthesized images, yet these object-rich conditions showed weaker alignment. This dissociation reveals that DNNs capture the texture-statistical component of early visual responses while failing to explain later, object-related variance.

Authors

Keywords

No keywords available for this article.