Latest AI and machine learning research in ophthalmology for healthcare professionals.
Imageomics uses machine learning to accelerate our understanding of biological traits and human disease processes. Some of the earliest imageomics applications used deep learning to assess human diseases. For example, retinal fundus images were analyzed to diagnose diabetic retinopathy. The imaging modality optical coherence tomography (OCT) is widely used to diagnose and monitor the progression o...
Vision-Language-Action (VLA) models rely on current observations, including images, language instructions, and robot states, to predict actions and complete tasks. While accurate visual perception is crucial for precise action prediction and execution, recent work has attempted to further improve performance by introducing explicit reasoning during inference. However, such approaches face signific...
Generalization to novel visual conditions remains a central challenge for both human and machine vision, yet standard robustness metrics offer limited...
Continuum robots possess high flexibility and redundancy, making them well suited for safe interaction in complex environments, yet their continuous d...
Recent 3D CT vision-language models align volumes with reports via contrastive pretraining, but typically rely on limited public data and provide only...
The classification of Intangible Cultural Heritage (ICH) images in the Mekong Delta poses unique challenges due to limited annotated data, high visual...
Safe autonomous systems in complex environments require robust road anomaly segmentation to identify unknown obstacles. However, existing approaches o...
We introduce V-SONAR, a vision-language embedding space extended from the text-only embedding space SONAR (Omnilingual Embeddings Team et al., 2026), ...
Clinically reliable perception of surgical scenes is essential for advancing intelligent, context-aware intraoperative assistance such as instrument h...
Foundation vision models are increasingly adopted in medical image analysis. Due to domain shift, these pretrained models misalign with medical image ...
Medical Vision-Language Models have shown promising potential in clinical decision support, yet they remain prone to factual hallucinations due to ins...
Large Vision-Language Models (LVLMs) have adopted visual token pruning strategies to mitigate substantial computational overhead incurred by extensive...
Reinforcement learning (RL) is increasingly used to post-train medical Vision-Language Models (VLMs), yet it remains unclear whether RL improves medic...
Transforming image features from perspective view (PV) space to bird's-eye-view (BEV) space remains challenging in autonomous driving due to depth amb...
Recent multimodal models such as Contrastive Language-Image Pre-training (CLIP) have shown remarkable ability to align visual and linguistic represent...
Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens increas...
Vision-language models (VLMs) show strong potential for complex diagnostic tasks in medical imaging. However, applying VLMs to multi-organ medical ima...
Fine-grained crop-weed segmentation is essential for enabling targeted herbicide application in precision agriculture. However, existing deep learning...
Vision-language-action (VLA) models integrate visual observations and language instructions to predict robot actions, demonstrating promising generali...
Optical design is the process of configuring optical elements to precisely manipulate light for high-fidelity imaging. It is inherently a highly non-c...