Large Language Models vs. NASS Common Coding Scenarios: An Analysis of Accuracy and Financial Implications using 2026 CPT AMA Codebook Guidelines.

Journal: The spine journal : official journal of the North American Spine Society
Published Date:

Abstract

BACKGROUND CONTEXT: Accurate Current Procedural Terminology (CPT) coding is essential for compliant revenue cycle management in spine surgery. However, increasing documentation burdens and the complexity of bundling rules often lead to significant revenue leakage or audit risk. PURPOSE: To evaluate the baseline efficacy of three state-of-the-art Large Language Models (LLMs) in generating accurate CPT codes compared to the "Gold Standard" North American Spine Society (NASS) Common Coding Scenarios, and to quantify the downstream financial impact. STUDY DESIGN/SETTING: Comparative artificial intelligence performance analysis. PATIENT SAMPLE: Twenty standardized clinical vignettes selected from the NASS Common Coding Scenarios guide, representing a broad spectrum of spine pathology including cervical decompression/fusion, lumbar decompression, lumbar fusion, and complex deformity. OUTCOME MEASURES: Exact match accuracy, error of omission (under-coding), error of commission (over-coding/unbundling), and net financial variance utilizing 2026 Medicare Physician Fee Schedule (MPFS) Work RVU (wRVU) values mapped to a standardized commercial conversion factor of $60/wRVU. METHODS: Each vignette was input into three LLMs: GPT-4o (OpenAI), Gemini 1.5 Pro (Google), and Claude 4.6 Sonnet (Anthropic). Under isolated zero-shot conditions, the models were prompted to function as certified professional coders and generate the appropriate CPT codes and modifiers. Outputs were scored against the official NASS answer key, and subsequent financial variances were calculated. RESULTS: Gemini 1.5 Pro achieved an exact match rate of 65%, followed by ChatGPT-4o (55%) and Claude 4.6 Sonnet (40%), though these differences in overall performance were not statistically significant (p = 0.28). However, performance varied substantially by procedure type; notably, both ChatGPT and Claude achieved a 0% exact match rate for Lumbar Fusion constructs, frequently failing to correctly bundle interbody and posterior instrumentation codes. Financially, Gemini 1.5 Pro was the most stable, resulting in a negligible mean net revenue variance of -$2.16 per case. In contrast, ChatGPT and Claude demonstrated a strong tendency toward over-coding and high financial volatility, resulting in mean net revenue overcharges of +$75.60 and +$104.07 per case, respectively. CONCLUSIONS: While no statistically significant difference in overall accuracy was detected between models, qualitative deficits remain severe in handling complex fusion hierarchies. This study evaluates baseline performance under isolated zero-shot conditions; the tendency of Claude and ChatGPT to create unbundled billable codes presents a severe audit risk, while Gemini demonstrated a lower mean financial variance. Currently, LLMs should function only as adjunctive tools requiring strict human supervision.

Authors

Keywords

No keywords available for this article.