Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3881-3900 of 9,853 articles

DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering

Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered significant attention due to its practical applications. In this paper, we introduce the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. Specifically, we propo...

Jul 27 2026 2607.23921v1

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI

Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence. This review provides an integrated overview of CFP AI through...

Jul 27 2026 2607.23972v1
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality imag...

Jul 27 2026 2607.24064v1
UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handle...

Jul 27 2026 2607.24157v1
Myopia Prevention and Control 3.0: Artificial Intelligence--Driven Risk Stratification, Proactive Monitoring, and Personalized Intervention

The convergence of artificial intelligence (AI), digital sensing, and ubiquitous computing has created an unprecedented opportunity to transform myopi...

Jul 27 2026 2607.24187v1
Surgical Re-enactment for Operating Room Workflow Datasets

The introduction of new technologies, such as surgical robots, is driving the vision of a connected, smart operating room (OR). However, realizing thi...

Jul 27 2026 2607.24206v1
MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features ...

Jul 27 2026 2607.24424v1
KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability

Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In cl...

Jul 27 2026 2607.24730v1
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundam...

Jul 27 2026 2607.24743v1
Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in...

Jul 26 2026 2607.23493v1
GNM Head: A Generative aNthropometric Model of the human head

Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstructio...

Jul 26 2026 2607.23687v1
PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morpholog...

Jul 26 2026 2607.23794v1
TextSLIP: Text Self-Supervised CLIP for Medical Report Generation

Automating radiology report generation is important for improving reporting consistency and clinical workflows . While Contrastive Language--Image Pre...

Jul 24 2026 2607.21970v1
LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR

End-to-end OCR systems based on vision-language models have achieved strong performance in complex document OCR, but their efficiency is limited by th...

Jul 24 2026 2607.22200v1
Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints

While Large Vision-Language Models (LVLMs) exhibit strong perceptual capabilities, they remain vulnerable in visual reasoning tasks. Existing benchmar...

Jul 23 2026 2607.21722v1
Sparse Concept Channels in Frozen 3D CT Vision Encoders

Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know which internal units enc...

Jul 23 2026 2607.20993v1
GeoThreat: Transferable Targeted Adversarial Attacks on Large Vision-Language Models for Remote Sensing Image Interpretation

Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing their robustness in cross-modal semantic und...

Jul 23 2026 2607.21036v1
Do Pathology Vision-Language Models Truly See Pathology?

Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. Ho...

Jul 23 2026 2607.21065v1
Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification

Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reduci...

Jul 23 2026 2607.21068v1
CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiri...

Jul 23 2026 2607.21155v1
Browse Categories