Impact of Labeling Inaccuracy and Image Noise on Tooth Segmentation in Panoramic Radiographs using Federated, Centralized and Local Learning.

Journal: Dento maxillo facial radiology
Published Date:

Abstract

OBJECTIVES: Federated learning (FL) may mitigate privacy constraints, heterogeneous data quality, and inconsistent labeling in dental diagnostic artificial intelligence (AI). FL was compared with centralized (CL) and local learning (LL) for tooth segmentation in panoramic radiographs across multiple data corruption scenarios. METHODS: An Attention U-Net was trained on 2066 radiographs from six institutions across four settings: baseline (unaltered data); label manipulation (dilated/missing annotations); image-quality manipulation (additive Gaussian noise); and exclusion of one faulty client with corrupted data. FL was implemented via the Flower AI framework. Per-client training- and validation loss trajectories were monitored for anomaly detection and a set of metrics (Dice, IoU, HD, HD95 and ASSD) were evaluated on a hold-out test set. From these metrics significance results were reported through Wilcoxon signed-rank test. CL and LL served as comparators. RESULTS: Baseline: FL achieved a median Dice of 0.94889 (ASSD: 1.33229), slightly better than CL at 0.94706 (ASSD: 1.37074) and LL at 0.93557-0.94026 (ASSD: 1.51910-1.69777). Label manipulation: FL maintained the best median Dice score at 0.94884 (ASSD: 1.46487) versus CL's 0.94183 (ASSD: 1.75738) and LL's 0.93003-0.94026 (ASSD: 1.51910-2.11462). Similar performance was observed when two faulty clients were introduced. Image noise: FL led with Dice at 0.94853 (ASSD: 1.31088); CL scored 0.94787 (ASSD: 1.36131); LL ranged from 0.93179-0.94026 (ASSD: 1.51910-1.77350). Similar performance was observed when two faulty clients were introduced, with CL performing slightly better than FL. Faulty-client exclusion: FL reached Dice at 0.94790 (ASSD: 1.33113) better than CL's 0.94550 (ASSD: 1.39318). Loss-curve monitoring reliably flagged the corrupted site. CONCLUSIONS: FL matches or exceeds CL and outperforms LL across corruption scenarios while preserving privacy. Per-client loss trajectories provide an effective anomaly-detection mechanism and support FL as a practical, privacy-preserving approach for scalable clinical AI deployment.

Authors

  • Johan Andreas Balle Rubak
    Department of Dentistry and Oral Health, Aarhus University, Vennelyst Boulevard 9, 8000, Aarhus, Denmark.
  • Khuram Naveed
    Department of Dentistry and Oral Health, Aarhus University, Vennelyst Boulevard 9, 8000, Aarhus, Denmark.
  • Sanyam Jain
    Bombay Hospital and Research Centre, Mumbai, India.
  • Lukas Esterle
    Department of Electrical- and Computer Engineering, Aarhus University, Finlandsgade 22, 8200, Aarhus, Denmark.
  • Alexandros Iosifidis
  • Ruben Pauwels
    Aarhus Institute of Advanced Studies (AIAS), Aarhus University, Høegh-Guldbergs Gade 6B, 8000-C, Aarhus, Denmark. [email protected].

Keywords

No keywords available for this article.