Latest AI and machine learning research in ophthalmology for healthcare professionals.
Retinal vessel segmentation is a vital early detection method for several severe ocular diseases. Despite significant progress in retinal vessel segmentation with the advancement of Neural Networks, there are still challenges to overcome. Specifically, retinal vessel segmentation aims to predict the class label for every pixel within a fundus image, with a primary focus on intra-image discrimina...
3D mask presentation attack detection is crucial for protecting face recognition systems against the rising threat of 3D mask attacks. While most existing methods utilize multimodal features or remote photoplethysmography (rPPG) signals to distinguish between real faces and 3D masks, they face significant challenges, such as the high costs associated with multimodal sensors and limited generaliz...
Panoramic imaging enables capturing 360{\deg} images with an ultra-wide Field-of-View (FoV) for dense omnidirectional perception. However, current p...
The growing demand for intelligent logistics, particularly fine-grained terminal delivery, underscores the need for autonomous UAV (Unmanned Aerial ...
We propose manvr3d, a novel VR-ready platform for interactive human-in-the-loop cell tracking. We utilize VR controllers and eye-tracking hardware t...
Medical AI assistants support doctors in disease diagnosis, medical image analysis, and report generation. However, they still face significant chal...
Introduction: Data from wearable devices collected in free-living settings, and labelled with physical activity behaviours compatible with health re...
Recent advances in multimodal Reward Models (RMs) have shown significant promise in delivering reward signals to align vision models with human pref...
Beamforming is a well-known technique to combine signals from multiple sensors. It has a wide range of application domains. This paper introduces th...
Real-world machine learning models require rigorous evaluation before deployment, especially in safety-critical domains like autonomous driving and ...
The Transformer architecture has achieved significant success in natural language processing, motivating its adaptation to computer vision tasks. Un...
Adversarial attacks have been fairly explored for computer vision and vision-language models. However, the avenue of adversarial attack for the visi...
Synthesizing interactive 3D scenes from text is essential for gaming, virtual reality, and embodied AI. However, existing methods face several chall...
Stroke is a major public health problem, affecting millions worldwide. Deep learning has recently demonstrated promise for enhancing the diagnosis a...
Learning discriminative 3D representations that generalize well to unknown testing categories is an emerging requirement for many real-world 3D appl...
Vision-language pretraining on large datasets of images-text pairs is one of the main building blocks of current Vision-Language Models. While with ...
Vision Foundation Models (VFMs) are large-scale, pre-trained models that serve as general-purpose backbones for various computer vision tasks. As VF...
Underwater visual enhancement (UVE) and underwater 3D reconstruction pose significant challenges in computer vision and AI-based tasks due to comp...
High-quality fundus images provide essential anatomical information for clinical screening and ophthalmic disease diagnosis. Yet, due to hardware li...
Surgical workflow recognition is vital for automating tasks, supporting decision-making, and training novice surgeons, ultimately improving patient ...