Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4421-4440 of 9,853 articles

The Limits of Learning from Pictures and Text: Vision-Language Models and Embodied Scene Understanding

What information is sufficient to learn the full richness of human scene understanding? The distributional hypothesis holds that the statistical co-occurrence of language and images captures the conceptual knowledge underlying visual cognition. Vision-language models (VLMs) are trained on massive paired text-image corpora but lack embodied experience, making them an ideal test of the distributiona...

Mar 27 2026 2603.26589v1

ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?

Beneath the stunning visual fidelity of modern AIGC models lies a "logical desert", where systems fail tasks that require physical, causal, or complex spatial reasoning. Current evaluations largely rely on superficial metrics or fragmented benchmarks, creating a ``performance mirage'' that overlooks the generative process. To address this, we introduce ViGoR Vision-G}nerative Reasoning-centric Ben...

Mar 26 2026 2603.25823v1
Can Vision Foundation Models Navigate? Zero-Shot Real-World Evaluation and Lessons Learned

Visual Navigation Models (VNMs) promise generalizable, robot navigation by learning from large-scale visual demonstrations. Despite growing real-world...

Mar 26 2026 2603.25937v1
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs

Zero-shot referring expression comprehension (REC) aims to locate target objects in images given natural language queries without relying on task-spec...

Mar 26 2026 2603.25004v1
Vision Hopfield Memory Networks

Recent vision and multimodal foundation backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable progress, ...

Mar 26 2026 2603.25157v1
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion

Purpose: The integration of multimodal imaging into operating rooms paves the way for comprehensive surgical scene understanding. In ophthalmic surger...

Mar 26 2026 2603.25555v1
LanteRn: Latent Visual Structured Reasoning

While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs). As a result, mos...

Mar 26 2026 2603.25629v1
MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models

Vision Foundation Models (VFMs) have become the cornerstone of modern computer vision, offering robust representations across a wide array of tasks. W...

Mar 26 2026 2603.25744v1
VOLMO: Versatile and Open Large Models for Ophthalmology

Vision impairment affects millions globally, and early detection is critical to preventing irreversible vision loss. Ophthalmology workflows require c...

Mar 25 2026 2603.23953v2
ReDiPrune: Relevance-Diversity Pre-Projection Token Pruning for Efficient Multimodal LLMs

Recent multimodal large language models are computationally expensive because Transformers must process a large number of visual tokens. We present \t...

Mar 25 2026 2603.24680v1
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration

Recent advances in multimodal vision-language models (VLMs) have enabled joint reasoning over visual and textual information, yet their application to...

Mar 25 2026 2603.24696v1
Gaze patterns predict preference and confidence in pairwise AI image evaluation

Preference learning methods, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on pairwise huma...

Mar 25 2026 2603.24849v1
Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding

Large Vision-Language Models (VLMs) have achieved remarkable success in multi-modal reasoning, but their inference time efficiency remains a significa...

Mar 25 2026 2603.23914v1
Revealing Multi-View Hallucination in Large Vision-Language Models

Large vision-language models (LVLMs) are increasingly being applied to multi-view image inputs captured from diverse viewpoints. However, despite this...

Mar 25 2026 2603.23934v1
VOLMO: Versatile and Open Large Models for Ophthalmology

Vision impairment affects millions globally, and early detection is critical to preventing irreversible vision loss. Ophthalmology workflows require c...

Mar 25 2026 2603.23953v1
Machine vision with small numbers of detected photons per inference

Machine vision, including object recognition and image reconstruction, is a central technology in many consumer devices and scientific instruments. Th...

Mar 25 2026 2603.23974v1
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification

Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barr...

Mar 25 2026 2603.24058v1
Retinal Layer Segmentation in OCT Images With 2.5D Cross-slice Feature Fusion Module for Glaucoma Assessment

For accurate glaucoma diagnosis and monitoring, reliable retinal layer segmentation in OCT images is essential. However, existing 2D segmentation meth...

Mar 25 2026 2603.24115v1
Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection

Current Large Vision Language Models (LVLMs) excel at many zero-shot tasks like image captioning, visual question answering and OCR. However, these sa...

Mar 25 2026 2603.24181v1
Connecting Meteorite Spectra to Lunar Surface Composition Using Hyperspectral Imaging and Machine Learning

We present an innovative, cost-effective framework integrating laboratory Hyperspectral Imaging (HSI) of the Bechar010 Lunar meteorite with ground-bas...

Mar 25 2026 2603.24323v1
Browse Categories