Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4661-4680 of 9,853 articles

Evaluating Graphical Perception Capabilities of Vision Transformers

Vision Transformers, ViTs, have emerged as a powerful alternative to convolutional neural networks, CNNs, in a variety of image-based tasks. While CNNs have previously been evaluated for their ability to perform graphical perception tasks, which are essential for interpreting visualizations, the perceptual capabilities of ViTs remain largely unexplored. In this work, we investigate the performance...

Feb 20 2026 2602.18178v1

Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models

Vision-language models (VLMs) have made substantial progress across a wide range of visual question answering benchmarks, spanning visual reasoning, document understanding, and multimodal dialogue. These improvements are evident in a wide range of VLMs built on a variety of base models, alignment architectures, and training data. However, recent works show that these models trail behind in traditi...

Feb 19 2026 2602.17871v1
Machine Learning Based Prediction of Surgical Outcomes in Chronic Rhinosinusitis from Clinical Data

Artificial intelligence (AI) has increasingly transformed medical prognostics by enabling rapid and accurate analysis across imaging and pathology. Ho...

Feb 19 2026 2602.17888v1
Selective Training for Large Vision Language Models via Visual Information Gain

Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying on...

Feb 19 2026 2602.17186v1
GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. This la...

Feb 19 2026 2602.17200v1
IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models

We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolve ...

Feb 18 2026 2602.16138v1
Breaking the Sub-Millimeter Barrier: Eyeframe Acquisition from Color Images

Eyeframe lens tracing is an important process in the optical industry that requires sub-millimeter precision to ensure proper lens fitting and optimal...

Feb 18 2026 2602.16281v1
VETime: Vision Enhanced Zero-Shot Time Series Anomaly Detection

Time-series anomaly detection (TSAD) requires identifying both immediate Point Anomalies and long-range Context Anomalies. However, existing foundatio...

Feb 18 2026 2602.16681v1
Are Object-Centric Representations Better At Compositional Generalization?

Compositional generalization, the ability to reason about novel combinations of familiar concepts, is fundamental to human cognition and a critical ch...

Feb 18 2026 2602.16689v1
Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation

Visual loco-manipulation of arbitrary objects in the wild with humanoid robots requires accurate end-effector (EE) control and a generalizable underst...

Feb 18 2026 2602.16705v1
Xray-Visual Models: Scaling Vision models on Industry Scale Data

We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data....

Feb 18 2026 2602.16918v1
A Large-Scale Computer-Vision Mapping of the Geometric Structures of Stroboscopically-Induced Visual Hallucinations

Visual hallucinations (VHs) occur across psychedelic states and diverse psychiatric and neurological conditions, yet their phenomenology remains diffi...

Local REM sleep-N1-wake sleep stage mixing in narcolepsy type 1

Type 1 narcolepsy (NT1), a disorder caused by the loss of hypocretin/orexin transmission, is characterized by daytime sleepiness and symptoms where Ra...

Visual Persuasion: What Influences Decisions of Vision-Language Models?

The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (VLMs). Th...

Feb 17 2026 2602.15278v1
Seeing to Generalize: How Visual Data Corrects Binding Shortcuts

Vision Language Models (VLMs) are designed to extend Large Language Models (LLMs) with visual capabilities, yet in this work we observe a surprising p...

Feb 16 2026 2602.15183v1
Disentangling physiological heterogeneity in retinal aging using a deep learning-based biological age framework

Biological age estimators quantify aging-related variation but provide limited insight into organ-specific aging processes. The retina enables non-inv...

TikArt: Aperture-Guided Observation for Fine-Grained Visual Reasoning via Reinforcement Learning

We address fine-grained visual reasoning in multimodal large language models (MLLMs), where key evidence may reside in tiny objects, cluttered regions...

Feb 16 2026 2602.14482v1
VIPA: Visual Informative Part Attention for Referring Image Segmentation

Referring Image Segmentation (RIS) aims to segment a target object described by a natural language expression. Existing methods have evolved by levera...

Feb 16 2026 2602.14788v1
ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery

Vision language models (VLMs) achieve strong performance on RGB imagery, but they do not generalize to thermal images. Thermal sensing plays a critica...

Feb 16 2026 2602.14989v1
MarsRetrieval: Benchmarking Vision-Language Models for Planetary-Scale Geospatial Retrieval on Mars

Data-driven approaches like deep learning are rapidly advancing planetary science, particularly in Mars exploration. Despite recent progress, most exi...

Feb 15 2026 2602.13961v1
Browse Categories