Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4181-4200 of 9,853 articles

Dual-Integrated Low-Latency Single-Lens Infrared Computational Imaging for Object Detection

Computational imaging enables compact infrared systems, but deep-learning pipelines that combine image reconstruction and object detection often introduce substantial inference latency. Most existing acceleration strategies compress the reconstruction network while overlooking physical priors from the optical path, leaving a trade-off between accuracy and speed. We present Physics-aware Dual-Integ...

May 21 2026 2605.21964v1

Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and recent work extends RLVR to multimodal large language models (MLLMs). This transfer, however, surfaces a faithfulness challenge: faithful perception of task-relevant visual evidence and faithful use of that evidence during reasoning, leading to uns...

May 21 2026 2605.22072v1
JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materi...

May 21 2026 2605.22080v1
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on ...

May 21 2026 2605.22089v1
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence

Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential...

May 21 2026 2605.22414v1
Matching with Deliberation: Test-Time Evolutionary Hierarchical Multi-Agents for Zero-Shot Compositional Image Retrieval

Zero-Shot Compositional Image Retrieval (ZS-CIR) requires both preserving the visual continuity of the reference image and faithfully executing the se...

May 21 2026 2605.22478v1
SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pr...

May 21 2026 2605.22536v1
Synthetic Data Alone is Enough? Rethinking Data Scarcity in Pediatric Rare Disease Recognition

Children with rare genetic diseases often exhibit distinctive facial phenotypes, yet developing computer vision systems for early diagnosis remains ch...

May 21 2026 2605.22767v1
MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data

Real-time cognitive load assessment from eye-tracking signals could potentially enable adaptive human-centered-AI such as safety-critical applications...

May 21 2026 2605.22775v1
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, exi...

May 21 2026 2605.22812v1
Language-dependent diagnostic safety of medical AI systems: a cross-lingual benchmarking and prospective clinical study

Background Patients worldwide receive healthcare in many languages, yet medical AI systems are validated almost exclusively in high-resource languages...

Fully Homomorphic Collaborative Learning for Safe Cross-Healthcare Institution Development and Implementation of Foundation Models

Foundation models (FMs) are powerful tools to allow the broad clinical application of artificial intelligence (AI) in healthcare systems, offering ada...

Predictive coding video models capture dorsal parietal representations and human judgments for surfaces defined by motion

Stimulus-computable models have transformed our understanding of ventral visual processing, yet comparable progress in modeling the dorsal visual stre...

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models

Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by...

May 18 2026 2605.17826v1
SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reaso...

May 18 2026 2605.17949v1
A More Word-like Image Tokenization for MLLMs

Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a seque...

May 18 2026 2605.17954v1
See Silhouettes in Motion with Neuromorphic Vision

Quasi-bimodal objects, such as text, road signs, and barcodes, play a basic yet vital role in daily visual communication. By boiling these down to cle...

May 18 2026 2605.17984v1
Semi-LAR: Semi-supervised Contrastive Learning with Linear Attention for Removal of Nighttime Flares

Lens flare removal is challenging due to the large spatial extent of flare artifacts and their entanglement with scene structures, while existing meth...

May 18 2026 2605.18156v1
Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology

Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational histopathology...

May 18 2026 2605.18419v1
What is Holding Back Latent Visual Reasoning?

Humans can approach complex visual problems by mentally simulating intermediate visual steps, rather than reasoning through language alone. Inspired b...

May 18 2026 2605.18445v1
Browse Categories