Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4001-4020 of 9,853 articles

GlaKG: A Biomarker-Centric Fundus Knowledge Graph for Explainable Glaucoma Diagnosis and Risk Assessment

Glaucoma is a leading cause of irreversible blindness worldwide, yet most automated diagnosis systems rely on opaque deep-learning models that offer little clinical interpretability. We present GlaKG, a biomarker-centric fundus knowledge graph that integrates structural biomarkers, clinically grounded rules, and image features to produce traceable reasoning for glaucoma diagnosis and risk stratifi...

Jul 6 2026 2607.04673v1

Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language Models

Vision-language models (VLMs) perform well on visual question answering with high-quality images but struggle when questions require knowledge beyond what is clearly and directly visible. In such settings, uncertainty quantification should not only indicate whether the model is likely to fail but also diagnose why it is uncertain, across dimensions such as perception, entity recognition, and knowl...

Jul 6 2026 2607.04683v1
Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions ...

Jul 6 2026 2607.05122v1
ChatImage: Navigating Long-Form LLM Answers through Interactive Images

Large Language Models (LLMs) can produce detailed answers to complex queries, but these answers are typically presented as dense linear text, which ma...

Jul 6 2026 2607.05290v1
InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics

Camera intrinsics are vital for recovering 3D structure from 2D video. However, most 3D algorithms assume fixed intrinsics throughout a video, an assu...

Jul 6 2026 2607.05389v1
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model

Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on ...

Jul 6 2026 2607.05396v1
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answ...

Jul 5 2026 2607.04163v1
IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation

While Large Vision-Language Models (VLMs) demonstrate remarkable generic capabilities, their clinical reasoning in specialized domains like ocular sur...

Jul 5 2026 2607.04344v1
Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach

In this work, we study the last-meter precision navigation for UAVs, e.g., autonomously reaching a target within the final 10 meters using monocular v...

Jul 5 2026 2607.04352v1
TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction

Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal r...

Jul 5 2026 2607.04484v1
Network-mediated diffusion produces disordered self-organization in vegetation

Vegetation patterns in arid and semiarid regions emerge as a result of a self-organization process triggered by water scarcity. While highly regular p...

Comparing Artificial Intelligence versus Human Screening in Systematic Reviews

Introduction. Systematic reviews are essential for informing health policy and practice. Artificial intelligence (AI) automates the article screening ...

Shared Pain, Shared Decisions: How Empathy Shapes Social Conformity Through Physiology and Visual Attention

Empathy enables individuals to attune to others' experiences through shared affective, sensorimotor, and neural representations, but its influence on ...

LLM-Empowered Multimodal Fusion Framework for Autonomous Driving: Semantic Enhancement and Channel-Adaptive Design

Vision-radar fusion is central to robust autonomous driving, combining dense visual semantics with precise range and velocity measurements from radar....

Jul 2 2026 2607.01772v1
Approximate Attention Weighting for Sustainable FPGA-Based Vision Transformer Inference

Vision Transformers have reshaped computer vision by using self-attention to capture global context across image regions. This makes them attractive f...

Jul 2 2026 2607.01798v1
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a new trend in...

Jul 2 2026 2607.01982v1
Show Me Examples: Inferring Visual Concepts from Image Sets

Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current ...

Jul 2 2026 2607.02402v1
Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning

Large vision-language models can reason over multimodal inputs by generating textual chains of thought (CoT). A key capability exhibited in CoT reason...

Jul 2 2026 2607.02490v1
Towards Robustness against Typographic Attack with Training-free Concept Localization

Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Model...

Jul 2 2026 2607.02494v1
Retinal resuscitation in post-mortem eyes

Vision loss compromises the quality of life of millions of people worldwide. Currently, vision-restoring therapies are lacking. Post-mortem preservati...

Browse Categories