Latest AI and machine learning research in ophthalmology for healthcare professionals.
Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLAs primarily focus on direct input--output mappings, lacking the intermediate reasoning steps cruci...
We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. JEEM includes the tasks of image captioning and visual question answering, and features culturally rich and regionally diverse content. This dataset aims to assess the ability of VLMs to generalize across dialec...
Workflows are a fundamental component of automation in enterprise platforms, enabling the orchestration of tasks, data processing, and system integr...
Visual in-context learning (VICL), as a new paradigm in computer vision, allows the model to rapidly adapt to various tasks with only a handful of p...
In this work, we aim to compress the vision tokens of a Large Vision Language Model (LVLM) into a representation that is simultaneously suitable for...
Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot cap...
Accurate camera calibration is a fundamental task for 3D perception, especially when dealing with real-world, in-the-wild environments where complex...
Our research is motivated by the urgent global issue of a large population affected by retinal diseases, which are evenly distributed but underserve...
Most multimodal large language models (MLLMs) treat visual tokens as "a sequence of text", integrating them with text tokens into a large language m...
Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonst...
Transformer-based models have driven significant advancements in Multimodal Large Language Models (MLLMs), yet their computational costs surge drast...
Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual ...
Weakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). R...
We introduce Vision as LoRA (VoRA), a novel paradigm for transforming an LLM into an MLLM. Unlike prevalent MLLM architectures that rely on external...
The rapid development of multimodal large language models has resulted in remarkable advancements in visual perception and understanding, consolidat...
Lung cancer remains one of the leading causes of cancer-related mortality worldwide. A crucial challenge for early diagnosis is differentiating unce...
Exposure-based interventions rely on inhibitory learning, often studied through Pavlovian conditioning. While disgust conditioning is increasingly l...
Multimodal large language models (MLLMs) have demonstrated significant potential in medical Visual Question Answering (VQA). Yet, they remain prone ...
Can Visual Language Models (VLMs) effectively capture human visual preferences? This work addresses this question by training VLMs to think about pr...
High-resolution perception of visual details is crucial for daily tasks. Current vision pre-training, however, is still limited to low resolutions (...