Stimulus-Driven Leakage in Naturalistic Neuroimaging
Journal:
bioRxiv
Published Date:
Mar 25, 2026
Abstract
This article elucidates a methodological pitfall of cross-validation for evaluating predictive models applied to naturalistic neuroimaging data---namely, "stimulus-driven leakage." While this problem has been well known as "leakage in training examples" in machine learning, it may be difficult to detect in practice due to conventions in neuroscience. Stimulus-driven leakage can occur when predictive modelling is applied to data from a conventional neuroscientific design, characterised by a limited set of stimuli repeated across trials and/or participants. It results in spurious predictive performance due to overfitting to repeated signals, even in the presence of independent noise. Through comprehensive simulations and real-world examples, following a theoretical formulation, the article underscores how such data leakage can occur and how severely it can compromise results and conclusions when combined with widely spread informal reverse inference. The article concludes with practical recommendations for researchers to avoid stimulus-driven leakage in their experimental design and analysis.