Latest AI and machine learning research in ophthalmology for healthcare professionals.
Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are limited by the coarse granularity of spatial control of the editing results. In contrast, visual prompts such as drag and point can provide precise spatial guidance, but are limited by the inherent ambiguity in semantic i...
Multimodal large language models (MLLMs) excel at visual reasoning but rely on text-based chain-of-thought (CoT), lacking interpretable visual intermediates. Existing methods use opaque tokens or external tools, missing key properties. We propose Gen-VCoT, a framework using expert vision models to generate RGB images as reasoning intermediates. It has three stages: visual grounding (SAM segmentati...
Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving t...
Vision-Language Models (VLMs) are AI systems that process both images and text, yet they often struggle with compositional visual reasoning questions ...
X-ray contraband detection is critical for security in large-scale logistics and transportation, yet conventional detectors struggle to adapt to emerg...
This work proposes an integrated pipeline for automatic glaucoma detection method from easily available colour fundas images based on an adaptive algo...
Spiking neural networks (SNNs) are brain-inspired, event-driven models that compute with sparse spikes, which enables highly efficient visual percepti...
Physical adversarial attacks on vision systems are typically studied through scene manipulation, such as adversarial patches or projections, where the...
Study Objectives To quantify high-resolution video-based eye kinematics across sleep macro- and microstructure and determine their coupling with pupil...
Community science platforms like iNaturalist generate unprecedented volumes of biodiversity data, but their scientific utility depends critically on a...
Vision-language models (VLMs) can answer image-based questions confidently, and often correctly, even when no image is provided. This mirage behavior ...
Background. Mice make substantial eye movements during head-fixed visual stimulation, and uncorrected gaze shifts corrupt receptive field measurements...
Visual Text Comprehension (VTC) renders text into images for a vision-language model (VLM) to read, sidestepping LLM context-window limits and powerin...
Vision-language models (VLMs) achieve strong singleshot spatial grounding, yet lack any mechanism to observe and correct their own predictions. We fin...
Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. ...
Background and Objective: Falls among elderly people can cause serious injury and reduce quality of life. Timely prediction and detection are essentia...
Large Vision-Language Models (LVLMs) have achieved strong performance across medical imaging tasks, yet they remain prone to factual inconsistencies, ...
Modern Vision-Language Models (VLMs) benefit from chain-of-thought prompting and test-time scaling, but these gains often come with prohibitive infere...
Understanding multi-label images remains a challenging task in computer vision. With the rapid progress of vision-language multimodal learning, vision...
Visual causal reasoning is essential for understanding and intervening in the physical world, requiring identification of causal variables from visual...