Large Language Model-Based Localization of Premature Ventricular Contraction Origins: A Retrospective Diagnostic Accuracy Study.
Journal:
Journal of cardiovascular electrophysiology
Published Date:
Aug 1, 2026
Abstract
BACKGROUND: Accurate localization of premature ventricular contraction (PVC) origin from 12-lead electrocardiography (ECG) is important for procedural planning in catheter ablation. Although convolutional neural network (CNN)-based models have shown promising diagnostic performance, they require task-specific training and remain limited in interpretability. We evaluated whether large language model (LLM)-based ECG image interpretation could perform binary left-versus-right PVC origin localization from 12-lead ECG images while providing a traceable diagnostic process. METHODS: We retrospectively studied 157 patients who underwent successful catheter ablation of PVCs or idiopathic ventricular tachycardia. ECG images were classified as RIGHT-origin (n = 103) or LEFT-origin (n = 54) according to the final successful ablation site. Three approaches were compared: a CNN-based baseline model, an LLM-based one-shot approach, and an LLM-based staged extraction framework with deterministic rule-based integration. Performance was evaluated across five independent seeds using positive predictive value (PPV), negative predictive value (NPV), recall, and PPV + NPV. RESULTS: The CNN-based baseline model achieved a PPV of 0.546 ± 0.058 and an NPV of 0.793 ± 0.047, with area under the curve values ranging from 0.595 to 0.730. The staged extraction framework generated a continuous rule-based score with discrimination numerically comparable to the CNN-based model (AUC 0.720 ± 0.045 vs. 0.712 ± 0.054). A stricter threshold, determined from training data, increased PPV at the expense of recall and warrants prospective external validation of this operating point. CONCLUSIONS: LLM-based staged extraction demonstrated the potential to achieve binary left-versus-right PVC origin localization from 12-lead ECG images while providing a traceable stepwise diagnostic process. The continuous rule-based score showed discrimination numerically comparable to the CNN-based model, and the strict threshold-determined using training data-identified a potential high-PPV operating point that requires prospective external validation.
Authors
Keywords
No keywords available for this article.