Artificial Intelligence Accurately Assists in Billing for Orthopaedic Lower Extremity Surgery: Performance of the Mistral-NeMo Language Model.
Journal:
Arthroscopy : the journal of arthroscopic & related surgery : official publication of the Arthroscopy Association of North America and the International Arthroscopy Association
Published Date:
Sep 6, 2026
Abstract
PURPOSE: To elucidate Mistral-NeMo's proficiency as a novel artificial intelligence assistant to verify coding accuracy and improve efficiency in manual medical coding practices in orthopaedic surgery. METHODS: This study tested Mistral-NeMo on 1000 operative notes labeled with the Current Procedural Terminology (CPT) codes from 177 providers. In total, there were 46 unique CPT codes; the most common were 29881 (knee arthroscopy with meniscectomy for both medial and lateral menisci), 29880 (knee arthroscopy with meniscectomy for 1 meniscus), and 29888 (ACL repair/reconstruction). Each model prompt included an operative note with either the true CPT code associated with the note or a randomly selected incorrect CPT code to serve as positive or negative controls. Trials asked the model for either a binary "Yes/No" response or a confidence score (0-100). RESULTS: The results from the binary-response trials showed that the Mistral-NeMo model correctly identified 90% (n = 1000) of the correct CPT codes for each operative note and rejected 99.80% (n = 1000) of incorrect codes (P < .001). The model achieved a precision of 99.80% and a recall of 90% under the testing framework. The model's performance was subanalyzed for the most common CPT codes, 29881, 29880, and 29888, accounting for 747 of the total operative notes. The model labeled the notes with these CPT codes with accuracy of 95.60% (n = 430), 95.50% (n = 160), and 82.50% (n = 157), respectively (P < .001). Among the confidence score trials, Mistral-NeMo showed an area under the operating curve of 0.96 and 0.97, indicating high classification ability. CONCLUSIONS: The Mistral-NeMo language model showed high accuracy in classifying CPT codes for femur- and knee-related surgical operative notes when CPT billing descriptions were provided. In contrast, model performance was insignificant in the absence of billing descriptions, indicating dependence on contextual information for accurate classification. CLINICAL RELEVANCE: This study assesses the performance of artificial intelligence, specifically the Mistral-NeMo language model, in the role of assisting in automating the billing process. Incidental coding error, along with the overall demands of billing, place a significant burden on both clinicians and administrative staff that divert their attention away from patient care. Investigation of artificial intelligence-based techniques enables automated validation of billing codes. Reducing coding errors has the potential to enhance clinical workflow and decrease the administrative burden on surgical practices.
Authors
Keywords
No keywords available for this article.