The Basic Science of Large Language Models in Orthopaedic Surgery.

Journal: The Journal of the American Academy of Orthopaedic Surgeons
Published Date:

Abstract

Orthopaedic surgeons routinely consult search engines, journals, and curated websites to stay current on orthopaedic knowledge. The emergence of large language models, such as OpenAI ChatGPT and Google MedGemma, is changing the way we search for information and how residents learn. Although many orthopaedic surgeons are users of artificial intelligence (AI), most are uncertain about how these tools actually work and why they sometimes give impressively accurate explanations alongside glaring factual errors and fabricated citations. This review provides an overview of the underlying preclinical studies behind large language models at the level of detail needed to empower orthopaedic surgeons with the knowledge needed to critically evaluate AI outputs, design future research projects, and effectively incorporate AI tools into clinical practice and resident education. Through clinical examples including a Schatzker VI tibial plateau fracture and an L4 pedicle screw sizing question, we illustrate two distinct classes of AI failure-retrieval failures and reasoning failures-and demonstrate how understanding the preclinical studies behind these errors equips surgeons to evaluate any AI tool regardless of where or how it runs.

Authors

Keywords

No keywords available for this article.