Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4641-4660 of 9,853 articles

Seeing Through Words: Controlling Visual Retrieval Quality with Language Models

Text-to-image retrieval is a fundamental task in vision-language learning, yet in real-world scenarios it is often challenged by short and underspecified user queries. Such queries are typically only one or two words long, rendering them semantically ambiguous, prone to collisions across diverse visual interpretations, and lacking explicit control over the quality of retrieved images. To address t...

Feb 24 2026 2602.21175v1

Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning

While Vision-Language Models (VLMs) exhibit exceptional 2D visual understanding, their ability to comprehend and reason about 3D space--a cornerstone of spatial intelligence--remains superficial. Current methodologies attempt to bridge this domain gap either by relying on explicit 3D modalities or by augmenting VLMs with partial, view-conditioned geometric priors. However, such approaches hinder s...

Feb 24 2026 2602.21186v1
Squint: Fast Visual Reinforcement Learning for Sim-to-Real Robotics

Visual reinforcement learning is appealing for robotics but expensive -- off-policy methods are sample-efficient yet slow; on-policy methods paralleli...

Feb 24 2026 2602.21203v1
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, curre...

Feb 23 2026 2602.19768v2
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device

Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry a...

Feb 23 2026 2602.20161v2
Decoupling Vision and Language: Codebook Anchored Visual Adaptation

Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders oft...

Feb 23 2026 2602.19449v1
RAID: Retrieval-Augmented Anomaly Detection

Unsupervised Anomaly Detection (UAD) aims to identify abnormal regions by establishing correspondences between test images and normal templates. Exist...

Feb 23 2026 2602.19611v1
Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness

Vision language models (VLMs) have achieved remarkable success in broad visual understanding, yet they remain challenged by object-centric reasoning o...

Feb 23 2026 2602.19615v1
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, curre...

Feb 23 2026 2602.19768v1
ApET: Approximation-Error Guided Token Compression for Efficient VLMs

Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibi...

Feb 23 2026 2602.19870v1
A Computationally Efficient Multidimensional Vision Transformer

Vision Transformers have achieved state-of-the-art performance in a wide range of computer vision tasks, but their practical deployment is limited b...

Feb 23 2026 2602.19982v1
HeatPrompt: Zero-Shot Vision-Language Modeling of Urban Heat Demand from Satellite Images

Accurate heat-demand maps play a crucial role in decarbonizing space heating, yet most municipalities lack detailed building-level data needed to calc...

Feb 23 2026 2602.20066v1
StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

Edge-based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We exte...

Feb 23 2026 2602.20089v1
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device

Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry a...

Feb 23 2026 2602.20161v1
StreetTree: A Large-Scale Global Benchmark for Fine-Grained Tree Species Classification

The fine-grained classification of street trees is a crucial task for urban planning, streetscape management, and the assessment of urban ecosystem se...

Feb 22 2026 2602.19123v1
Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing

With the rapid development of remote sensing image archives, asking questions about images has become an effective way of gathering specific informati...

Feb 22 2026 2602.19217v1
WildOS: Open-Vocabulary Object Search in the Wild

Autonomous navigation in complex, unstructured outdoor environments requires robots to operate over long ranges without prior maps and limited depth s...

Feb 22 2026 2602.19308v1
RetinaVision: XAI-Driven Augmented Regulation for Precise Retinal Disease Classification using deep learning framework

Early and accurate classification of retinal diseases is critical to counter vision loss and for guiding clinical management of retinal diseases. In t...

Feb 22 2026 2602.19324v1
Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation Spaces

The Platonic Representation Hypothesis posits that learned representations from models trained on different modalities converge to a shared latent str...

Feb 22 2026 2602.19367v1
ZACH-ViT: Regime-Dependent Inductive Bias in Compact Vision Transformers for Medical Imaging

Vision Transformers rely on positional embeddings and class tokens that encode fixed spatial priors. While effective for natural images, these priors ...

Feb 20 2026 2602.17929v1
Browse Categories