Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4301-4320 of 9,853 articles

Modulating Cross-Modal Convergence with Single-Stimulus, Intra-Modal Dispersion

Neural networks exhibit a remarkable degree of representational convergence across diverse architectures, training objectives, and even data modalities. This convergence is predictive of alignment with brain representation. A recent hypothesis suggests this arises from learning the underlying structure in the environment in similar ways. However, it is unclear how individual stimuli elicit converg...

Apr 23 2026 2604.21836v1

Directional Confusions Reveal Divergent Inductive Biases Through Rate-Distortion Geometry in Human and Machine Vision

Humans and modern vision models can reach similar classification accuracy while making systematically different kinds of mistakes - differing not in how often they err, but in who gets mistaken for whom, and in which direction. We show that these directional confusions reveal distinct inductive biases that are invisible to accuracy alone. Using matched human and deep vision model responses on a na...

Apr 23 2026 2604.21909v1
Validating a Deep Learning Algorithm to Identify Patients with Glaucoma using Systemic Electronic Health Records

We evaluated whether a glaucoma risk assessment (GRA) model trained on All of Us national data can identify patients at high probability of glaucoma u...

Apr 22 2026 2604.20921v1
Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback

Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. How...

Apr 22 2026 2604.20730v2
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry

Vision evaluations are typically done through multi-step processes. In most contemporary fields, experts analyze images using structured, evidence-bas...

Apr 22 2026 2604.20983v1
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models

Vision-language models benefit from high-resolution images, but the increase in visual-token count incurs high compute overhead. Humans resolve this t...

Apr 22 2026 2604.21079v1
Weighting What Matters: Boosting Sample Efficiency in Medical Report Generation via Token Reweighting

Training vision-language models (VLMs) for medical report generation is often hindered by the scarcity of high-quality annotated data. This work evalu...

Apr 22 2026 2604.21082v1
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval

Composed image retrieval, multi-turn composed image retrieval, and composed video retrieval all share a common paradigm: composing the reference visua...

Apr 22 2026 2604.20318v1
Image Generators are Generalist Vision Learners

Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent ...

Apr 22 2026 2604.20329v1
X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely...

Apr 22 2026 2604.20350v1
Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models

Object Referring-guided Scanpath Prediction (ORSP) aims to predict the human attention scanpath when they search for a specific target object in a vis...

Apr 22 2026 2604.20361v1
Evian: Towards Explainable Visual Instruction-tuning Data Auditing

The efficacy of Large Vision-Language Models (LVLMs) is critically dependent on the quality of their training data, requiring a precise balance betwee...

Apr 22 2026 2604.20544v1
Beyond ZOH: Advanced Discretization Strategies for Vision Mamba

Vision Mamba, as a state space model (SSM), employs a zero-order hold (ZOH) discretization, which assumes that input signals remain constant between s...

Apr 22 2026 2604.20606v1
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

Reinforcement learning (RL) with verifiable rewards (RLVR) has demonstrated the great potential of enhancing the reasoning abilities in multimodal lar...

Apr 22 2026 2604.20705v1
Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback

Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. How...

Apr 22 2026 2604.20730v1
Seeing Candidates at Scale: Multimodal LLMs for Visual Political Communication on Instagram

This paper presents a computational case study that evaluates the capabilities of specialized machine learning models and emerging multimodal large la...

Apr 21 2026 2604.19489v1
Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning

Assessing chronic wound infection from photographs is challenging because visual appearance varies across wound etiologies, anatomical locations, and ...

Apr 21 2026 2604.19937v1
DistortBench: Benchmarking Vision Language Models on Image Distortion Identification

Vision-language models (VLMs) are increasingly used in settings where sensitivity to low-level image degradations matters, including content moderatio...

Apr 21 2026 2604.19966v1
Subject-Aware Multi-Granularity Alignment for Zero-Shot EEG-to-Image Retrieval

Zero-shot EEG-to-image retrieval aims to decode perceived visual content from electroencephalography (EEG) by aligning neural responses with pretraine...

Apr 20 2026 2604.17782v1
OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models

Vision-Language Models(VLMs) excel at autoregressive text generation, yet end-to-end autonomous driving requires multi-task learning with structured o...

Apr 20 2026 2604.17915v1
Browse Categories