Exploring the role of reinforcement learning in vision-language models for cardiovascular disease decision support.
Journal:
Journal of biomedical informatics
Published Date:
Apr 1, 2026
Abstract
OBJECTIVE: To explore the role of reinforcement learning (RL) in vision-language models (VLMs) for cardiovascular disease (CVD) decision support and assess whether RL-enhanced multimodal reasoning improves clinical classification performance and interpretability across datasets. METHODS: We propose CVD-thinking, a multimodal framework that integrates electrocardiogram (ECG) images with textual reports, demographic information, and laboratory results from MIMIC-IV and MIMIC-IV-ECG. The model is trained using Group Relative Policy Optimization (GRPO) with Clinical Classification Reward (CCR), a generalizable reward template that combines asymmetric outcome utilities with a soft F1 shaping term to mitigate reward sparsity. We use explicit reasoning prompts () to elicit step-by-step rationales, compare against supervised fine-tuning (SFT) and a GPT-4o baseline, and evaluate external validity on an independent Mayo Clinic dataset. Ablation studies assess the contributions of reward shaping and explicit reasoning. RESULTS: On the internal test set, GRPO training improved precision compared with baseline approaches (CVD-thinking: 0.716 vs. SFT: 0.662 vs. GPT-4o: 0.625). Explicit reasoning further improved performance relative to a non-reasoning variant (0.716 vs. 0.598) while producing traceable rationales. These gains generalized to the external Mayo Clinic dataset. Ablations suggest that reward shaping and explicit reasoning both contribute to performance and reasoning-format adherence. CONCLUSION: RL can improve the reliability and transparency of multimodal reasoning for CVD clinical decision support, but its impact on performance depends on the reward formulation. Model scale also matters: smaller backbones (e.g., 3B parameters) show less stable optimization and can experience performance degradation.
Authors
Keywords
No keywords available for this article.