Stress or Arousal? Exercise Confounding in Wearable Stress Detection.

Journal: Medical engineering & physics
Published Date:

Abstract

Wearable stress-detection models are commonly validated by separating rest from stress, but this does not establish that their outputs are specific to stress rather than broader physiological activation. We re-analysed two public wearable datasets and made a matched within-dataset stress-specificity experiment the primary test. PhysioNet event tags were used to reconstruct protocol-defined rest and stress-induction blocks; 29 participants with matched aerobic and anaerobic recordings contributed 522 rest, 192 stress, and 3431 exercise-session windows. Four feature-based classifiers were trained to distinguish rest from stress under nested leave-one-participant-out evaluation. Preprocessing, hyperparameter selection, and probability-threshold calibration used training participants only, and exercise-session data were excluded from all model development. The best calibrated model, XGBoost, achieved balanced accuracy of 0.703, stress detection of 0.604, and rest specificity of 0.801, yet labelled 82.6% of held-out exercise-session windows as stress. At participant level, its exercise-session false-stress rate was 0.840 [0.775, 0.897], compared with a rest false-stress rate of 0.221 [0.151, 0.301], giving a paired difference of 0.619 [0.537, 0.699]. Protocol-defined stress blocks also had higher self-reported stress than rest blocks (mean difference 1.22 [0.89, 1.53] on the supplied 1-10 scale; 27/29 participants; one-sided Wilcoxon p=4.64×10⁻⁶). Adding time-domain heart-rate-variability features did not materially change the result. A Wearable Stress and Affect Detection (WESAD)-to-PhysioNet transfer analysis remained poor but was treated as secondary evidence because it also contains dataset and protocol shift. These findings show that exercise-associated activation can produce substantial false-stress responses even when stress recognition and an exercise-session challenge are evaluated within the same dataset and the challenge data are excluded from model development.

Authors

Keywords

No keywords available for this article.