Latest AI and machine learning research in ophthalmology for healthcare professionals.
Vision-language-action (VLA) models extend vision-language models (VLM) by integrating action generation modules for robotic manipulation. Leveraging strengths of VLM in vision perception and instruction understanding, VLA models exhibit promising generalization across diverse manipulation tasks. However, applications demanding high precision and accuracy reveal performance gaps without further ...
Vision Language Models (VLMs) hold great promise for streamlining labour-intensive medical imaging workflows, yet systematic security evaluations in clinical settings remain scarce. We introduce VSF--Med, an end-to-end vulnerability-scoring framework for medical VLMs that unites three novel components: (i) a rich library of sophisticated text-prompt attack templates targeting emerging threat vec...
Navigating everyday social situations often requires juggling conflicting goals, such as conveying a harsh truth, maintaining trust, all while still...
Accurate distance estimation is a fundamental challenge in robotic perception, particularly in omnidirectional imaging, where traditional geometric ...
The aim of this work is to obtain precise atmospheric parameters and chemical abundances automatically for solar twins and analogs to find signature...
Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebo...
Recent advancements in multimodal large language models have enhanced document understanding by integrating textual and visual information. However,...
OBJECTIVES: To compare the effectiveness of expert-designed machine learning models and code-free automated machine learning (AutoML) models in classi...
Recent Multimodal Large Language Models (MLLMs) excel on benchmark vision-language tasks, yet little is known about how input visual quality shapes ...
Recent advancements in medical image analysis have led to the development of highly specialized models tailored to specific clinical tasks. These mo...
Previous methods for image geo-localization have typically treated the task as either classification or retrieval, often relying on black-box decisi...
This work addresses image restoration tasks through the lens of inverse problems using unpaired datasets. In contrast to traditional approaches -- w...
Recent advancements in vision-language systems have improved the accuracy of Radiological Visual Question Answering (VQA) Models. However, some chal...
Large Vision-Language Models (VLMs) have demonstrated potential in enhancing mobile robot navigation in human-centric environments by understanding ...
Data-driven approaches struggle with precise manipulation; imitation learning requires many hard-to-obtain demonstrations, while reinforcement learn...
Even during fixation the human eye is constantly in low amplitude motion, jittering over small angles in random directions at up to 100Hz. This moti...
With the growing integration of vision-language models (VLMs), mobile agents are now widely used for tasks like UI automation and camera-based user ...
Although Large Vision Language Models (LVLMs) have demonstrated remarkable performance in image understanding tasks, their computational efficiency ...
Person re-identification (ReID) has evolved from handcrafted feature-based methods to deep learning approaches and, more recently, to models incorpo...
Despite the widespread adoption of transformers in medical applications, the exploration of multi-scale learning through transformers remains limite...