Comparative Analysis of Agreement and Scoring Patterns Between Human Evaluators and Artificial Intelligence: The Case of Complete Denture Tooth Arrangement.
Journal:
European journal of dental education : official journal of the Association for Dental Education in Europe
Published Date:
Jul 17, 2026
Abstract
OBJECTIVE: This study aimed to compare agreement and scoring patterns between human evaluators and artificial intelligence (AI) models in the rubric-based evaluation of complete denture design applications. METHOD: Thirty complete denture fabrication assignments prepared by students in the Dental Prosthetics Technology Programme were photographed from five different angles according to a standard protocol and evaluated using a 20-item rubric. The assignments were independently assessed by three human evaluators with at least 5 years of experience in complete dentures, as well as by the AI models Gemini 3 Pro, Claude Sonnet 4.6, and Grok 4. Inter-group agreement was analysed using Fleiss' Kappa, while scoring differences were analysed using the Friedman and Wilcoxon Signed-Rank Tests. RESULTS: While significant agreement was observed among human evaluators (κ = 0.63; p < 0.001), agreement among the AI models remained low (κ = 0.09; p = 0.012). AI models assigned significantly higher scores than human evaluators (1.72 ± 0.50 vs. 1.09 ± 0.85; p < 0.001). CONCLUSION: Current AI models do not yet demonstrate sufficient reliability to function as independent evaluators of laboratory-based tasks requiring precise visual and spatial assessment. Consequently, final assessment evaluation should therefore remain under the responsibility of experienced dental educators, with AI tools used only as a complementary feedback tool.
Authors
Keywords
No keywords available for this article.