EndoVLM: A Vision-Language Assistant for Gastrointestinal Endoscopy.
Journal:
Journal of imaging informatics in medicine
Published Date:
Aug 26, 2026
Abstract
Gastrointestinal endoscopy generates extensive high-resolution video data, posing significant challenges for efficient and accurate computer-aided diagnosis of gastrointestinal diseases. To address this, we propose EndoVLM (Endoscopy Vision-Language Model), a specialized visual question-answering assistant for gastroenterology. EndoVLM introduces ConvNeXt as a hierarchical visual encoder to replace traditional ViTs (Vision Transformers), inherently compressing high-resolution gastrointestinal endoscopy images into information-dense features while reducing redundant token overhead. In addition, EndoVLM adopts a three-stage fine-tuning schedule that progressively aligns the visual-language projector, adapts the ConvNeXt backbone to endoscopic imagery, and improves instruction following in the language decoder. We position this design as an efficient adaptation strategy for high-resolution gastrointestinal endoscopy rather than as a new fine-tuning paradigm, with the aim of balancing token efficiency, domain adaptation, and deployment practicality in multimodal medical AI (artificial intelligence). Quantitatively, EndoVLM achieves 0.7326 ± 0.0038 ROUGE-1 (Recall-Oriented Understudy for Gisting Evaluation-1) and 0.5103 ± 0.0051 BLEU (Bilingual Evaluation Understudy) on Kvasir-VQA (Kvasir visual question answering), improving over the strongest public baseline by +0.0166 ROUGE-1 and +0.0323 BLEU, while the final three-stage model improves Gastrovision external validation from 0.4760 to 0.5799 Accuracy and from 0.1109 to 0.1268 macro-F1 (macro-averaged F1 score) over the representative two-stage baseline. The experimental section separates descriptive benchmark comparison from controlled follow-up analyses, including stage-wise convergence evidence, external validation on Gastrovision, and stage-wise training-strategy analysis.
Authors
Keywords
No keywords available for this article.