Deep Learning-Based Detection of Radiolucent Jaw Lesions on Panoramic Radiographs: A Comparative Study.
Journal:
Journal of stomatology, oral and maxillofacial surgery
Published Date:
Oct 1, 2026
Abstract
This retrospective study aimed to compare the performance of different deep learning-based object detection models for detecting radiolucent jaw lesions on panoramic radiographs. A total of 600 panoramic radiographs showing radiolucent jaw lesions were included in the study. All images were anonymized and manually annotated with rectangular bounding boxes using Roboflow. The dataset was divided into training, validation, and test subsets at a ratio of 80%, 10%, and 10%, respectively. Data augmentation was applied only to the training set, increasing the total number of images from 600 to 1,142. Four object detection models, YOLOv5, YOLOv8, YOLO26, and Faster R-CNN, were trained and evaluated using the same dataset split and experimental framework. Model performance was assessed using precision, recall, F1 score, and mean Average Precision at [email protected] and [email protected]:0.95. YOLOv8 achieved the highest precision value, with 79% precision, 59% recall, and a 68% F1 score. YOLOv5 achieved 65% precision, 76% recall, and a 70% F1 score; YOLO26 achieved 67% precision, 73% recall, and a 70% F1 score; and Faster R-CNN achieved 65% precision, 76% recall, and a 70% F1 score. Mean Average Precision at an IoU threshold of 0.50 was 60% for YOLOv8, 60% for YOLOv5, 67% for YOLO26 and 70% for Faster R-CNN, with corresponding [email protected]:0.95 values of 30%, 30%, 30% and 32%, respectively. Although the evaluated models showed generally comparable F1 scores, differences were observed in their precision-recall balance. Considering the potential speed disadvantage of two-stage detectors such as Faster R-CNN, YOLO-based models may represent more practical alternatives for clinical decision-support applications when comparable performance is obtained. The inclusion of radiolucent lesions from the entire panoramic field may have increased task complexity and contributed to the moderate performance values observed.
Authors
Keywords
No keywords available for this article.