Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

Journal: arXiv
Published Date:

Abstract

The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.

Authors

  • Yumi Lee; Harim Oh; Hyoryung Kim; Minji Kim; Eunsu Kim; Hyeseong Lee; Junya Fukuoka; Andrey Bychkov; Jijgee Munkhdelger; Rajiv Kumar Kaushal; Ayushi Sahay; Rajni Yadav; Bharathi Prabakaran; Sulen Sarioglu; Serdar Balcı; Ilknur Turkmen; Yuri Tolkach; Christian Harder; Julian Westerdorf; Reinhard Buettner; Audun Ljone Henriksen; Sepp De Raedt; Byung Hyun Lee; Sungjin Lim; Joohoon Lee; Gwanghyun Kim; Se Young Chun; Suryakant Singh; Saarthak Kapse; Prateek Prasanna; Kyung A Kim; Yousun Kang; Sehwan Yoo; Sungman Hong; Shubham Innani; Michael Feldman; Spyridon Bakas; Ujjwal Baid; Prasad Dutande; Suhas Gajare; Bhakti Baheti; Serkan Sökmen; Ece Tuğba Cebeci; Ahmet Halıcı; Musa Balcı; Kardelen Peçenek; Srividhya Sainath; Kyongseok Jang; Messi H. J. Lee; Noorul Wahab; Bodong Du; Jiaming Zhang; Qixiang Zhang; Jang-Hwan Choi; Sangjeong Ahn