Quantifying performance inflation from hidden sample dependence in biomedical image classification benchmarks
Journal:
bioRxiv
Published Date:
Oct 8, 2026
Abstract
Machine learning benchmarks often treat image files as independent observations, although multiple files may represent repeated views or other related samples. Prior work has shown that non-independent partitioning can inflate model performance, but the sensitivity of a biomedical benchmark to progressively increasing train-test relatedness is less often quantified. Here, we quantified how apparent classification performance changes when recoverable sample groups are progressively allowed to cross training, validation, and test partitions. Using source structure and filenames, we reconstructed analysis groups in a microscopy corpus compiled from multiple sources and generated five class stratified partitioning conditions, ranging from an intact group reference to random assignment of individual files. We applied the same training recipe within each model family to 129,173 images spanning 14 classes of blood cells and hematopoietic cell states. For a pretrained Vision Transformer, random assignment of individual files increased the mean test macro F1 from 0.853 to 0.977 and accuracy from 0.851 to 0.990 relative to the intact group reference. Apparent performance increased progressively across the intermediate conditions. The endpoint pattern also appeared with a ResNet-50 backbone, and the largest apparent gains occurred in immature granulocytic and blast classes. A descriptive analysis within the random assignment condition further showed higher scores among test images that had another member of the same group in the training set, although the exposed and unexposed subsets were not exchangeable. The grouping workflow, exact split definitions, code, and compact result files are released as HSD-Bench to support reproducible evaluation of hidden sample dependence.