Multi-Axial Analysis of Clinical Reasoning in Large Language Models: Inter-Verifier Disagreement and Its Implications for Automated Evaluation.

Journal: Journal of medical systems
Published Date:

Abstract

Evaluating clinical reasoning in large language models (LLMs) poses two open challenges: reference-oriented semantic metrics do not directly assess whether a model's stated diagnosis is supported by the evidence in its own justification, and the increasingly popular LLM-as-judge approach rests on a largely untested assumption-that independent verifier LLMs agree with one another. We assess three generator LLMs (HuatuoGPT-o1-8B, Meta-Llama-3.1-8B-Instruct, Meta-Llama-3.3-70B-Instruct) on 1,000 MIMIC-IV hospital-stay cases along four complementary axes (medical concept grounding, semantic similarity, semantic uncertainty, and evidence-conclusion coherence), with coherence judged independently by three frontier verifiers (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.4 mini). Two findings emerge. First, coherence reveals a dissociation that reference-oriented metrics do not capture: a model can score well on those axes yet still produce rationales that do not support its own conclusions. Second, inter-verifier agreement on coherence is consistently low (Fleiss' κ 0.087-0.223; disagreement 62.2%-74.3%), so the same rationale can be judged supported or unsupported depending on the verifier. A preliminary validation in which a physician adjudicated 50 cases echoed this: agreement with the physician varied across verifiers, underscoring that no single LLM reliably stands in for clinical assessment. Together, these results suggest a single LLM verifier lacks sufficient reliability to serve as a stand-alone judge of clinical reasoning at scale, and that structured human oversight remains essential. The unanimous-agreement tier offers a candidate for selective automation, but its clinical reliability remains to be confirmed in larger, multi-clinician adjudication studies.

Authors

Keywords

No keywords available for this article.