Alignment Between Large Language Models and Consensus in Endoscopic Spinal Disc Surgery: A Comparative Analysis Against the Neurocore-SENSED.
Journal:
JOR spine
Published Date:
Sep 4, 2026
Abstract
BACKGROUND: The Neurocore-SENSED framework, derived from a three-round modified Delphi process involving 77 international spine surgeons, provides a structured reference standard for reporting in endoscopic spine surgery (ESS) for disc disease. The ability of large language models (LLMs) to reproduce graded levels of expert agreement within a reporting framework has not been examined in ESS. OBJECTIVE: To evaluate the extent to which three contemporary LLMs align with the Neurocore-SENSED consensus and whether they discriminate between items of differing consensus level. METHODS: In January 2026, each of the 166 framework items with published item-level endorsement data was submitted once to GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.6, in independent sessions with web retrieval disabled. Models selected one of five ordered response options directly. Alignment was assessed by Spearman correlation with panel endorsement and by four measures of categorical agreement, each with bootstrap 95% confidence intervals. Separate protocols examined response stability, response format, and prior exposure to the consensus. RESULTS: All three models correlated significantly with panel endorsement. Gemini 3 Pro aligned most closely (r s = 0.709, 95% CI 0.627-0.776), exceeding Claude Sonnet 4.6 (r s = 0.622; p = 0.009) and GPT-5.2 (r s = 0.565; p < 0.001), which did not differ. Categorical agreement was limited and equivalent across models, with observed agreement of 0.506-0.518 and overlapping confidence intervals for all marginal-robust statistics. Between 81.3% and 95.8% of responses fell in the top two categories, and discordance was almost entirely unidirectional: 31 items were endorsed by every model but by fewer than half the panel, against one item in the opposite direction. Weighted kappa was not interpretable in this setting. CONCLUSION: Current LLMs reproduce the broad rank ordering of expert endorsement in ESS reporting but cannot reliably distinguish endorsed from contested propositions, principally because of a positive response bias.
Authors
Keywords
No keywords available for this article.