Image-Quality-Aware Multimodal Artificial Intelligence for Automated Structured OCT Report Generation in Glaucoma Evaluation.

Journal: Ophthalmology science
Published Date:

Abstract

OBJECTIVE: To develop an explainable multimodal large language model (MM-LLM) that (1) screens optic nerve head (ONH) OCT circle scans for quality and (2) generates structured clinical reports that include glaucoma diagnosis and sector-wise retinal nerve fiber layer (RNFL) thinning assessments. DESIGN: A retrospective cohort study using longitudinal data from the Diagnostic Innovations in Glaucoma Study and the African Descent and Glaucoma Evaluation Study. PARTICIPANTS: A total of 43 849 Spectralis circumpapillary B-scans centered on the ONH from 1310 subjects, including 1331 glaucomatous and 867 healthy eyes. METHODS: An MM-LLM (Llama 3.2 Vision-Instruct model) was fine-tuned to generate clinical descriptions of OCT imaging data. Training data included paired OCT images and automatically generated, structured clinical reports that described global and sectoral RNFL thinning. Poor-quality scans were labeled as unusable and paired with a fixed refusal statement. The model was evaluated on a held-out test set for 3 tasks: quality assessment, glaucoma detection, and RNFL thinning classification across 7 anatomical sectors. Evaluation metrics included accuracy, sensitivity, specificity, precision, and F1-score. Model description quality was also evaluated using standard text evaluation metrics (BLEU, ROUGE, METEOR, and BERTScore). MAIN OUTCOME MEASURES: Diagnostic accuracy metrics for each task; text evaluation metrics for description quality. RESULTS: The model achieved 0.90 accuracy and 0.98 specificity for quality triage. For glaucoma detection, accuracy was 0.86 (sensitivity 0.93, specificity 0.65, and F1-score 0.91). Retinal nerve fiber layer thinning prediction accuracy ranged from 0.83 to 0.94, with the highest performance in global, temporal, temporal superior, and temporal inferior sectors. Text generation scores (mean ± standard deviation) showed strong alignment with reference reports (BLEU: 0.82 ± 0.19; ROUGE-1: 0.94 ± 0.08; ROUGE-2: 0.87 ± 0.17; ROUGE-L: 0.92 ± 0.11; BERTScore-F1: 0.99 ± 0.02). Stratified analysis revealed better RNFL thinning detection in moderate-to-advanced glaucoma cases, especially in temporal sectors, while performance in nasal regions was better for mild cases. CONCLUSIONS: The fine-tuned MM-LLM generated accurate clinical descriptions based on OCT imaging. The model achieved high accuracy in identifying image quality issues and detecting glaucoma. The model provided sectoral descriptions of RNFL thinning to support clinical OCT evaluation. This approach shows potential as a scalable tool for clinical decision support, but further validation across additional datasets is needed. FINANCIAL DISCLOSURES: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Authors

Keywords

No keywords available for this article.