Special article-EchoPeer: a standardized framework for assessing echocardiography reports in the era of artificial intelligence: a recommendation from the Korean Society of Echocardiography AI and Future Strategy Committee.

Journal: Journal of cardiovascular imaging
Published Date:

Abstract

Artificial intelligence (AI) is rapidly advancing from automated measurement to full-report generation, yet existing frameworks do not provide a unified scoring approach for both human- and AI-authored reports. We developed EchoPeer, a three-step evaluation framework for echocardiography reports, scored against a reference standard. Step 0 (safety score, %) is a pass/fail safety gate for life-threatening findings (critical omission and hallucination). Step 1 (precision score, 0-80 scale) scores per-item diagnostic accuracy across 25 items in four anatomical domains. Step 2 (quality score, 0-20 scale) grades clinical utility across three checkpoints: relevance, synthesis, and clarity. The precision and quality scores are summed as a composite score (0-100 scale), with step 0 failures scored as 0. EchoPeer was applied to 30 adult transthoracic echocardiography cases interpreted by 11 human readers across three training levels. All three EchoPeer scores increased monotonically with training level. Median safety scores were 86.7%, 90.0%, and 96.7% for levels 1, 2, and 3, respectively (P = 0.005), and sensitivity for critical finding detection rose from 59.3% to 92.6% (P = 0.007). Median precision scores were 61.8, 64.7, and 70.9, respectively (P = 0.007), and median quality scores were 10.4, 12.9, and 15.1, respectively (P = 0.009). The composite score followed the same gradient (72.2, 77.7, and 86.2, respectively; P = 0.003), and the precision and quality scores correlated significantly at both the case level (Spearman ρ = 0.568) and the participant level (ρ = 0.888). EchoPeer discriminates clinical competence in human readers and produces an error profile that mirrors known challenges in echocardiographic reporting, providing a clinically grounded foundation for the future evaluation of AI-generated echocardiography reports.

Authors

Keywords

No keywords available for this article.