Artificial intelligence and vascular surgeons in patient communication: A comparative analysis of intelligibility and clinical appropriateness in acute deep vein thrombosis.

Journal: Phlebology
Published Date:

Abstract

ObjectiveEffective patient communication is critical in acute deep vein thrombosis (DVT) management. This study evaluated and compared the intelligibility and clinical appropriateness of patient-directed explanations for acute DVT generated by vascular surgeons versus a diverse panel of large language models (LLMs), using a balanced, unit-matched design.MethodsIn this observational, cross-sectional study, 11 board-certified vascular surgeons and 11 distinct, independently developed LLMs each produced one response to a standardized acute DVT vignette in Spanish. Responses were blinded and scored using an 8-item clinical appropriateness grid (two independent graders, consensus) and five validated Spanish readability indices. Central tendency was assessed with Mann-Whitney U tests; dispersion was assessed descriptively (coefficient of variation, range) and, exploratorily, with Brown-Forsythe tests.ResultsClinical appropriateness did not differ significantly between groups (median 7/8 in both; p = 0.258), dispersion was descriptively far smaller for LLMs (SD 0.40 vs 1.92; coefficient of variation 5.9% vs 32.5%). LLM responses were more readable on three of four content-complexity indices (Fernández-Huerta, Gutiérrez de Polini, Szigriszt-Pazos; all p < 0.02) and required a numerically lower educational grade level (Crawford 5.4 vs 6.1; p = 0.075). Readability dispersion was descriptively far lower for LLMs (Fernández-Huerta range 31 vs 200 points; coefficient of variation 15% vs 181%), driven in the surgeon cohort by a single response that fell far outside the interpretable range; formal variance tests did not reach significance at this sample size. LLM responses were longer (median 903 vs 354 words; p = 0.003) and took longer to read (4.5 vs 1.8 min; p < 0.001). Neither group addressed screening for silent pulmonary embolism. Free-versus-premium comparisons were model-specific rather than uniform.ConclusionCurrent-generation LLMs matched board-certified surgeons in clinical completeness and produced more readable, descriptively far more consistent patient-facing explanations, without the communication failures observed in a minority of human responses. Both cohorts share an important, guideline-adjacent counseling gap.

Authors

Keywords

No keywords available for this article.