Latest AI and machine learning research in ophthalmology for healthcare professionals.
This paper discusses how ophthalmologists often rely on multimodal data to improve diagnostic accuracy. However, complete multimodal data is rare in real-world applications due to a lack of medical equipment and concerns about data privacy. Traditional deep learning methods typically address these issues by learning representations in latent space. However, the paper highlights two key limitatio...
Palm-sized autonomous nano-drones, i.e., sub-50g in weight, recently entered the drone racing scenario, where they are tasked to avoid obstacles and navigate as fast as possible through gates. However, in contrast with their bigger counterparts, i.e., kg-scale drones, nano-drones expose three orders of magnitude less onboard memory and compute power, demanding more efficient and lightweight visi...
Recent advances in human preference alignment have significantly enhanced multimodal generation and understanding. A key approach is training reward...
STEAM education integrates Science, Technology, Engineering, Arts, and Mathematics to foster creativity and problem-solving. However, students with ...
Trajectory forecasting has become a popular deep learning task due to its relevance for scenario simulation for autonomous driving. Specifically, tr...
Vision-Language Models (VLMs) have emerged as powerful tools in artificial intelli-gence, capable of integrating textual and visual data for a unifi...
True intelligence hinges on the ability to uncover and leverage hidden causal relations. Despite significant progress in AI and computer vision (CV)...
Iris texture is widely regarded as a gold standard biometric modality for authentication and identification. The demand for robust iris recognition ...
Remote Sensing Image Captioning (RSIC) is a cross-modal field bridging vision and language, aimed at automatically generating natural language descr...
Visual servoing enables robots to precisely position their end-effector relative to a target object. While classical methods rely on hand-crafted fe...
The analysis of high-dimensional timeline data and the identification of outliers and anomalies is critical across diverse domains, including sensor...
We present Vinci, a vision-language system designed to provide real-time, comprehensive AI assistance on portable devices. At its core, Vinci levera...
Can our brain signals faithfully reflect the original visual stimuli, even including high-frequency details? Although human perceptual and cognitive...
Recently, Multimodal Large Language Models (MLLMs) have gained significant attention for their remarkable ability to process and analyze non-textual...
Automated defect detection in industrial manufacturing is essential for maintaining product quality and minimizing production errors. In air disc br...
Large multimodal models (LMMs) "see" images by leveraging the attention mechanism between text and visual tokens in the transformer decoder. Ideally...
LiDAR semantic segmentation models are typically trained from random initialization as universal pre-training is hindered by the lack of large, dive...
Accurate motion understanding of the dynamic objects within the scene in bird's-eye-view (BEV) is critical to ensure a reliable obstacle avoidance s...
Visual Language Models (VLMs) have demonstrated impressive capabilities in visual grounding tasks. However, their effectiveness in the medical domai...
3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its su...