CARE: Multimodal Alignment Learning for Ophthalmic Imaging via Joint Optimal-Transport Soft Matching and Graph Neural ODE Refinement.
Journal:
Journal of imaging informatics in medicine
Published Date:
Aug 31, 2026
Abstract
Medical artificial intelligence (AI) has shown great potential for the early screening and accurate diagnosis of fundus diseases. However, diagnostic methods based on single-modality retinal images still have clear limitations. Two-dimensional (2D) fundus images, including color fundus photography (CFP) and near-infrared (NIR) fundus images, capture superficial retinal morphology, vascular distribution, and lesion appearance, but they are insufficient for revealing deep microstructural changes. In contrast, optical coherence tomography (OCT) can characterize retinal layered structures and pathological deformations, yet lacks a comprehensive description of the overall fundus appearance. As a result, single-modality methods are often constrained by incomplete information in diagnosing complex retinal diseases. To exploit complementary information across modalities, multimodal approaches integrating fundus images and OCT have attracted increasing attention. However, learning reliable soft correspondence distributions between two-dimensional fundus images and OCT regions remains challenging. Due to substantial differences in imaging principles, spatial representations, and semantic granularity, complex many-to-many nonlinear correspondences commonly exist between the two modalities. Existing implicit attention-based fusion methods generally construct cross-modal interactions through independently normalized attention distributions, without explicitly imposing bilateral marginal constraints on the correspondence between fundus and OCT tokens. This may lead to unbalanced token utilization and modality dominance under noise interference or modality conflict, thereby limiting the stability of multimodal fusion. To address these issues, we propose a multimodal fundus disease diagnosis method based on progressive optimization of probabilistic correspondences. Specifically, entropy-regularized optimal transport is used to construct soft correspondence distributions between fundus image and OCT regions, while a unified graph structure and graph neural ordinary differential equation iteratively refine cross-modal alignment. In addition, modality-coupled contrastive learning is introduced to enhance disease discrimination and grading consistency. Experiments on the GAMMA and OLIVES datasets show that CARE achieves improved accuracy and Cohen's kappa under public benchmark settings, and demonstrates robustness under the evaluated noise perturbation and limited-data scenarios. Quantitatively, CARE achieved ACC/Kappa values of 0.860/0.893 on the GAMMA dataset and 1.000/1.000 on the OLIVES dataset. Statistical comparisons supported the main performance differences under the corresponding evaluation settings.
Authors
Keywords
No keywords available for this article.