Individual identification of dairy cows with occluded camera views using open-set contrastive learning model.
Journal:
Journal of dairy science
Published Date:
Jun 8, 2026
Abstract
This study aims to perform individual identification of dairy cows in freestall barns using computer vision techniques under challenging visual conditions such as partial occlusions, varying lighting, and diverse cow poses. Initially, cow identification was performed using the You Only Look Once, Version 8 model for detecting and annotating the animals, combined with radio frequency identification information. A dataset of 500 images per cow per day was collected for 49 cows over 4 d. A leave-one-day-out cross-validation strategy was applied for model evaluation, where data from one day were reserved for testing and the remaining 3 d were used for training and validation. Two deep learning approaches were explored for individual cow identification: the Xception model, which treats the problem as a multiclass classification task, and a contrastive learning approach implemented using a Siamese architecture to learn similarity between image pairs. In the Xception approach, each cow was treated as a distinct class. The contrastive learning model compared test images with reference embeddings generated from the training set, using k-means clustering to select representative samples per cow and assigning the identity based on the highest similarity. For the Xception model, the average precision across all test days was 0.84, recall was 0.79, F1-score was 0.78, and overall accuracy was 0.79. The best performance achieved by the Xception model across the test days was a precision of 0.89, recall of 0.88, F1-score of 0.87, and accuracy of 0.88. The contrastive learning model achieved an average precision of 0.61, recall of 0.70, F1-score of 0.64, and accuracy of 0.70 across all test days. Whereas its overall performance was lower than that of the Xception model in the closed-set scenario, the contrastive learning model was implemented in an open-set setting, allowing it to identify individuals not seen during training, and therefore demonstrated stronger generalization capabilities. When evaluated in an open-set setting, where images of cows not seen during training were included, it achieved an accuracy of 1.00, successfully identifying newly introduced animals in the herd as "unknown." This highlights its potential applicability in real-world environments with dynamic herd compositions, in which animals are moved in and out of specific groups. These findings highlight the strengths of both models: The Xception model achieved higher accuracy in the closed-set condition, whereas the contrastive learning model proved more effective in the open-set scenario, offering greater robustness to unseen individuals. In summary, computer vision can be seen as a potential alternative for traditional methods of animal identification, improving efficiency in monitoring livestock identity, reducing labor costs, minimizing human error, and allowing continuous data collection to support precision livestock farming and better decision-making. In addition, computer vision enables automated, scalable, and noninvasive phenotyping in conjunction with animal identification and tracking.
Authors
Keywords
No keywords available for this article.