Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3721-3740 of 9,853 articles

Automated 2D and 3D Segmentation of AMD and DME Lesions in OCT

Age-related macular degeneration (AMD) and diabetic macular edema (DME) are leading causes of vision loss, and optical coherence tomography (OCT) is the standard modality for detecting and monitoring the subtle lesions that drive treatment decisions. Most deep-learning segmentation work for OCT is validated only in-domain, leaving generalization to clinical data collected under different acquisiti...

Aug 27 2026 2608.27095v1

ReViCo: Unveiling the Limitations of VLMs in Visual Text Understanding via Error Correction

Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images. In this paper, we introduce ReViCo (Real Visual Correction), a benchmark designed to evaluate VLM text understanding through a novel task of visual text error correction. ReViCo challenges models to identify and fix text errors in real-world images, which...

Aug 27 2026 2608.27154v1
Vision-centric generative AI models: A software-hardware perspective

Vision generative artificial intelligence (AI) has emerged as one of the most rapidly advancing areas of deep learning. The explosion of multimodal mo...

Aug 27 2026 2608.27199v1
Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet t...

Aug 27 2026 2608.27417v1
What Do Medical Vision-Language Models Learn in Radiology? Transfer, Alignment, and Source-Proxy Leakage Under Distribution Shift

Medical vision-language models (VLMs) can appear reliable in-domain while failing when acquisition domain, paired supervision, or evaluation protocol ...

Aug 26 2026 2608.25251v1
V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

Vision-language-action (VLA) models provide a scalable path toward generalist robotic manipulation by integrating visual perception, language understa...

Aug 26 2026 2608.25308v1
GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models

Generative vision-language models (VLMs) are increasingly used in human-centered settings, yet they can produce demographically biased outputs even wh...

Aug 26 2026 2608.25375v1
Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

Autoregressive (AR) image generation has shown strong potential for scalable high-fidelity synthesis by modeling images as discrete token sequences. H...

Aug 26 2026 2608.25386v1
DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors

Self-supervised learning (SSL) encoders are vulnerable to backdoor attacks, posing threats to both visual SSL encoders and vision-language encoders. E...

Aug 26 2026 2608.25851v1
Visual General Intelligence: A White Paper

This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning ...

Aug 26 2026 2608.25924v1
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be...

Aug 26 2026 2608.26105v1
CHIASM: A Self-Supervised Visual Field Encoder for Neuro-Ophthalmology

Background: Artificial intelligence (AI) systems for glaucoma diagnosis and prognostication from visual fields (VF) are under active development, yet ...

Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking

Multi-camera 3D perception systems for warehouse scenes are trained largely on synthetic data and evaluated on physically captured environments. The r...

Aug 25 2026 2608.24130v1
X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis

Imaging factor disentanglement in text-to-image generation aims to independently control image acquisition properties such as types of camera lenses, ...

Aug 25 2026 2608.24563v1
Interpretable Fundus Image Classification via Ring-Based Retinal Vasculature Features

Retinal fundus photography is widely used for screening and monitoring ocular diseases, but many modern classification pipelines rely on deep latent r...

Aug 25 2026 2608.24723v1
What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal ...

Aug 24 2026 2608.23474v2
DriftAD: Visually-Guided Text Drift for Few-Shot Industrial Anomaly Detection

Few-shot anomaly detection (FSAD) has recently benefited from vision-language models such as CLIP, which enable anomaly de?tection by aligning visual ...

Aug 24 2026 2608.23723v1
Primate vision reveals a missing principle for robust dynamic AI

How does an intelligent visual system combine what objects look like with how they move while remaining robust as appearance changes? We addressed thi...

Aug 24 2026 2608.23790v1
LUX: A Lesion-Aware Graph-Conditioned Visual - Language Architecture for Explainable Endoscopic Captioning

The interpretation of endoscopic imagery in ulcerative colitis is complex and subjective, with variability in human assessment and subtle mucosal infl...

Aug 24 2026 2608.23853v1
HAP: Head-Adaptive Visual Token Pruning via Cross-Modal Alignment

Recent Vision-Language Models encode high-resolution images into long visual token sequences, incurring prohibitive prefill costs. To compress them, e...

Aug 24 2026 2608.23921v1
Browse Categories