Latest AI and machine learning research in ophthalmology for healthcare professionals.
Many everyday mobile manipulation tasks require precise interaction with small objects, such as grasping a knob to open a cabinet or pressing a light switch. In this paper, we develop Servoing with Vision Models (SVM), a closed-loop training-free framework that enables a mobile manipulator to tackle such precise tasks involving the manipulation of small objects. SVM employs an RGB-D wrist camera...
A large-scale vision and language model that has been pretrained on massive data encodes visual and linguistic prior, which makes it easier to generate images and language that are more natural and realistic. Despite this, there is still a significant domain gap between the modalities of vision and language, especially when training data is scarce in few-shot settings, where only very limited da...
Recent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting i...
We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational ...
Background: Mutations in KMT2B are a recognized cause of early-onset complex dystonia, with deep brain stimulation (DBS) of the internal globus pall...
Efficient evaluation of three-dimensional (3D) medical images is crucial for diagnostic and therapeutic practices in healthcare. Recent years have s...
Myopia, projected to affect 50% population globally by 2050, is a leading cause of vision loss. Eyes with pathological myopia exhibit distinctive sh...
The emergence of large Vision Language Models (VLMs) has broadened the scope and capabilities of single-modal Large Language Models (LLMs) by integr...
Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages. Consequently, evaluations of large vision-language mo...
State space models (SSMs) have recently garnered significant attention in computer vision. However, due to the unique characteristics of image data,...
Generalization remains a significant challenge for low-level vision models, which often struggle with unseen degradations in real-world scenarios de...
AI safety is a rapidly growing area of research that seeks to prevent the harm and misuse of frontier AI technology, particularly with respect to ge...
Optical Coherence Tomography (OCT) provides high-resolution cross-sectional images useful for diagnosing various diseases, but their distinct charac...
Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders the...
Visually linking matching cues is a crucial ability in daily life, such as identifying the same person in multiple photos based on their cues, even ...
Image-to-point cloud cross-modal Visual Place Recognition (VPR) is a challenging task where the query is an RGB image, and the database samples are ...
Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language model...
Accurate blur estimation is essential for high-performance imaging across various applications. Blur is typically represented by the point spread fu...
The rapid rise of AI-generated content has made detecting disinformation increasingly challenging. In particular, multimodal disinformation, i.e., o...
Humans experience the world through multiple modalities, such as, vision, language, and speech, making it natural to explore the commonality and dis...