Can Large Language Models Preserve Diagnostic Accuracy Despite Patient Self-Diagnosis and Framing Bias in Hypothetical Upper-Extremity Scenarios?

Journal: The Journal of hand surgery
Published Date:

Abstract

PURPOSE: Online health queries are often addressed by large language models (LLMs) embedded in search engines. It is possible that LLMs, like human clinicians, might be misdirected by vague symptom descriptions or inaccurate self-diagnoses. We examined patient and scenario factors associated with an LLM's ability to identify intended upper-extremity musculoskeletal diagnoses and its tendency to deviate from patient self-diagnoses in structured clinical vignettes. METHODS: ChatGPT (GPT-5) evaluated 180 randomized hypothetical clinical vignettes depicting five common upper-extremity conditions: de Quervain tendinopathy, rotator cuff tendinopathy, lateral epicondylitis, trigger digit, and trapeziometacarpal arthritis. Each vignette included randomized patient characteristics, characteristic or vague symptom descriptions, and a patient self-diagnosis (categorized as correct, a plausible alternative, or a common misconception diagnosis). The LLM was prompted to select the single most likely diagnosis. Multivariable logistic regression identified independent predictors of diagnostic accuracy and deviation. RESULTS: The LLM correctly identified the intended diagnosis in 165 of 180 scenarios (92%). Accuracy was unaffected by the patient's self-diagnosis, was higher for characteristic than vague symptom, and was lower for de Quervain tendinopathy relative to other conditions. The model deviated from the patient's proposed diagnosis in 117 scenarios (65%), of which 104 deviations (89%) appropriately aligned with the intended diagnosis. The LLM was more likely to disregard patient-provided diagnoses that did not match the intended diagnosis, regardless of whether they represented plausible alternatives or common misconceptions. CONCLUSIONS: In this experimental setting, an LLM identified simulated upper extremity conditions regardless of patient self-diagnosis, suggesting limited susceptibility to the anchoring, confirmation, and acquiescence biases known to affect human diagnostic reasoning. LLMs may therefore support debiasing and patient guidance by helping address unhealthy misconceptions and aligning tests and treatment choices with patient values. TYPE OF STUDY/LEVEL OF EVIDENCE: V (Experimental Vignette Diagnostic Accuracy Study).

Authors

Keywords

No keywords available for this article.