Latest AI and machine learning research in ophthalmology for healthcare professionals.
Glaucoma is a leading cause of irreversible blindness worldwide, yet most automated diagnosis systems rely on opaque deep-learning models that offer little clinical interpretability. We present GlaKG, a biomarker-centric fundus knowledge graph that integrates structural biomarkers, clinically grounded rules, and image features to produce traceable reasoning for glaucoma diagnosis and risk stratifi...
Vision-language models (VLMs) perform well on visual question answering with high-quality images but struggle when questions require knowledge beyond what is clearly and directly visible. In such settings, uncertainty quantification should not only indicate whether the model is likely to fail but also diagnose why it is uncertain, across dimensions such as perception, entity recognition, and knowl...
Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions ...
Large Language Models (LLMs) can produce detailed answers to complex queries, but these answers are typically presented as dense linear text, which ma...
Camera intrinsics are vital for recovering 3D structure from 2D video. However, most 3D algorithms assume fixed intrinsics throughout a video, an assu...
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on ...
Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answ...
While Large Vision-Language Models (VLMs) demonstrate remarkable generic capabilities, their clinical reasoning in specialized domains like ocular sur...
In this work, we study the last-meter precision navigation for UAVs, e.g., autonomously reaching a target within the final 10 meters using monocular v...
Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal r...
Vegetation patterns in arid and semiarid regions emerge as a result of a self-organization process triggered by water scarcity. While highly regular p...
Introduction. Systematic reviews are essential for informing health policy and practice. Artificial intelligence (AI) automates the article screening ...
Empathy enables individuals to attune to others' experiences through shared affective, sensorimotor, and neural representations, but its influence on ...
Vision-radar fusion is central to robust autonomous driving, combining dense visual semantics with precise range and velocity measurements from radar....
Vision Transformers have reshaped computer vision by using self-attention to capture global context across image regions. This makes them attractive f...
Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a new trend in...
Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current ...
Large vision-language models can reason over multimodal inputs by generating textual chains of thought (CoT). A key capability exhibited in CoT reason...
Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Model...
Vision loss compromises the quality of life of millions of people worldwide. Currently, vision-restoring therapies are lacking. Post-mortem preservati...