Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4261-4280 of 9,853 articles

Gated Subspace Inference for Transformer Acceleration

A method is presented for accelerating inference in transformer language models by exploiting the low effective rank of the token activation manifold at each layer. The method decomposes each activation vector into a subspace component and a residual, computes the linear-layer output on the subspace component via a cached low-rank weight image at reduced memory bandwidth, and applies a per-token g...

May 4 2026 2605.03109v1

Full-Field Stimulus Test for Visual Function Assessment in Ultra-Low Vision with Retinitis Pigmentosa

Purpose: Assessing visual function in patients with ultra-low vision (ULV), particularly those with retinitis pigmentosa (RP), remains a significant challenge in therapeutic development. Full-field stimulus test (FST) provides a quantitative measure of retinal light sensitivity and may serve as a valuable clinical endpoint. We investigated FST in ULV RP by examining its associations with functiona...

Jailbreaking Vision-Language Models Through the Visual Modality

The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak at...

May 1 2026 2605.00583v1
Deep Kernel Learning for Stratifying Glaucoma Trajectories

Effectively stratifying patient risk in chronic diseases like glaucoma is a major clinical challenge. Clinicians need tools to identify patients at hi...

May 1 2026 2605.00708v1
Modeling Subjective Urban Perception with Human Gaze

Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computatio...

May 1 2026 2605.00764v1
Let ViT Speak: Generative Language-Image Pre-training

In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining fra...

May 1 2026 2605.00809v1
The Effects of Visual Priming on Cooperative Behavior in Vision-Language Models

As Vision-Language Models (VLMs) become increasingly integrated into decision-making systems, it is essential to understand how visual inputs influenc...

Apr 30 2026 2604.27953v1
CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation

Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current v...

Apr 29 2026 2604.26288v2
ViBE: Visual-to-M/EEG Brain Encoding via Spatio-Temporal VAE and Distribution-Aligned Projection

Brain encoding models not only serve to decipher how visual stimuli are transformed into neural responses, but also represent a critical step toward v...

Apr 29 2026 2604.26218v1
CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation

Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current v...

Apr 29 2026 2604.26288v1
Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning

Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language l...

Apr 29 2026 2604.26370v1
Delineating Knowledge Boundaries for Honest Large Vision-Language Models

Large Vision-Language Models (VLMs) have achieved remarkable multimodal performance yet remain prone to factual hallucinations, particularly in long-t...

Apr 29 2026 2604.26419v1
Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners

One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack...

Apr 29 2026 2604.26488v1
Delta Score Matters! Spatial Adaptive Multi Guidance in Diffusion Models

Diffusion models have achieved remarkable success in synthesizing complex static and temporal visuals, a breakthrough largely driven by Classifier-Fre...

Apr 29 2026 2604.26503v1
3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis method for Single-View Aerial-Ground Re-Identification

Aerial-Ground Re-Identification (AG-ReID) is constrained by the viewpoint-domain gap, as drastic viewpoint disparities occlude or distort discriminati...

Apr 29 2026 2604.26520v1
AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision

Despite the rapid progress in data-driven 3D vision, aerial geometric 3D vision remains a formidable challenge due to the severe scarcity of large-sca...

Apr 29 2026 2604.26567v1
One Size Fits All? Comparing Foundation and Task-specific Models for Retinal Fluid Segmentation

Retinal fluids, detectable through optical coherence tomography (OCT), are key biomarkers for retinal diseases such as diabetic macular edema and age-...

Development and validation of a lesion-supervised deep learning system for diabetic retinopathy grading according to UK national screening criteria

Background: Diabetic retinopathy (DR) is the leading cause of preventable blindness among working-age adults worldwide, yet screening coverage remains...

Toward Multimodal Conversational AI for Age-Related Macular Degeneration

Despite strong performance of deep learning models in retinal disease detection, most systems produce static predictions without clinical reasoning or...

Apr 28 2026 2604.25720v1
Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning

Vision-Language Models (VLMs) exhibit strong performance in instruction following and open-ended vision-language reasoning, yet they frequently genera...

Apr 28 2026 2604.25809v1
Browse Categories