Latest AI and machine learning research in ophthalmology for healthcare professionals.
Recent advancements in multi-modal models have significantly improved vision-language alignment in radiology. However, existing approaches struggle to effectively utilize complex radiology reports for learning, rely on low-resolution images, and offer limited interpretability in attention mechanisms. To address these challenges, we introduce RadZero, a novel similarity-based cross-attention fram...
Medical image segmentation has achieved remarkable success through the continuous advancement of UNet-based and Transformer-based foundation backbones. However, clinical diagnosis in the real world often requires integrating domain knowledge, especially textual information. Conducting multimodal learning involves visual and text modalities shown as a solution, but collecting paired vision-langua...
This study presents a vision-guided robotic control system for automated fruit tree pruning applications. Traditional agricultural practices rely on...
Visual emotion analysis or recognition has gained considerable attention due to the growing interest in understanding how images can convey rich sem...
The practical application of AI tools for specific computer vision tasks relies on the "small-data regime" of hundreds to thousands of labeled sampl...
As automation and mobile robotics reshape work environments, rising expectations for productivity increase cognitive demands on human operators, lea...
Neurons in sensory systems encode stimulus information into their stochastic spiking response. The Mutual information has been broadly applied to th...
We present a full-spectrum machine learning framework for refractive index sensing using simulated absorption spectra from meta-grating structures c...
Chronic wounds affect a large population, particularly the elderly and diabetic patients, who often exhibit limited mobility and co-existing health ...
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual-text processing. However, existi...
Inspired by human visual attention, deep neural networks have widely adopted attention mechanisms to learn locally discriminative attributes for cha...
When a vision-language model (VLM) is prompted to identify an entity depicted in an image, it may answer 'I see a conifer,' rather than the specific...
Vision-language pre-training has recently gained popularity as it allows learning rich feature representations using large-scale data sources. This ...
The rapid evolution towards the sixth-generation (6G) networks demands advanced beamforming techniques to address challenges in dynamic, high-mobili...
Machine learning-based embedded systems for safety-critical applications, such as aerospace and autonomous driving, must be robust to perturbations ...
Recent progress in VLMs has demonstrated impressive capabilities across a variety of tasks in the natural image domain. Motivated by these advanceme...
We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks...
Typographic attacks exploit the interplay between text and visual content in multimodal foundation models, causing misclassifications when misleadin...
Mamba-based vision models have gained extensive attention as a result of being computationally more efficient than attention-based models. However, ...
Compositionality, or correctly recognizing scenes as compositions of atomic visual concepts, remains difficult for multimodal large language models ...