Recording context and data splitting can inflate acoustic re-identification accuracy across species and studies
Journal:
bioRxiv
Published Date:
Oct 9, 2026
Abstract
1. Re-identifying individual animals (re-ID) from their calls with machine learning (ML) could extend individual-based monitoring to species and areas hard to survey by capture and marking, and published studies often report re-ID accuracy above 90%. Two methodological practices can inflate these values: (1) attributing calls to individuals by territory, nest, tag or station, tying each animal to its own recording conditions, and (2) randomly splitting individual clips into training and test sets, allowing recording context to be shared between them. Each effect has so far been documented only in single-species studies. 2. We implement four diagnostics of how methodological choices affect re-ID accuracy: (1) splitting clips by recording session; (2) background tests using ambient recordings without the animal; (3) within-recording permutation and paired-clip tests where animals share recordings; and (4) open-set tests with thresholds set on other individuals. We applied them to 13 public datasets from nine species (eight birds, one bat), using BirdNET, 19 other neural networks trained on animal or general sound and 16 trained on human speech, each with one classifier. 3. Under session splits, BirdNET re-ID accuracy ranged from 0.16 to 0.97. Within one year, ambient sound alone carried substantial identity information for territorial chiffchaffs (BirdNET balanced accuracy 0.71, against 0.79 for song); background identification exceeded chance for 35 of 38 embeddings and controls. Count-matched random clip splits raised mean accuracy by 0.20 to 0.51 in 18 comparisons across three species. Where several animals shared recordings, only one of three datasets showed clearly that re-ID rested on the animal rather than the recording. With unfamiliar animals present and thresholds calibrated to target 10% stranger acceptance, across 36 neural networks and two scoring methods on wild session-split datasets, the highest median share of known birds' calls accepted and correctly named was 46%, using nearest-clip scoring with four enrolled birds; test-stranger acceptance could exceed the target. No neural network beat BirdNET on every dataset, and network rankings depended on layer choice. 4. The diagnostics need no new models and apply wherever a dataset holds what they require: session labels, background recordings, shared recordings or enough individuals. Re-ID accuracy reported without them can substantially overstate what acoustic re-ID achieves, and should not be used to plan population monitoring until they are applied.