Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5501-5520 of 9,853 articles

RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Radiology with Zero-Shot Multi-Task Capability

Recent advancements in multi-modal models have significantly improved vision-language alignment in radiology. However, existing approaches struggle to effectively utilize complex radiology reports for learning, rely on low-resolution images, and offer limited interpretability in attention mechanisms. To address these challenges, we introduce RadZero, a novel similarity-based cross-attention fram...

Zeus: Zero-shot LLM Instruction for Union Segmentation in Multimodal Medical Imaging

Medical image segmentation has achieved remarkable success through the continuous advancement of UNet-based and Transformer-based foundation backbones. However, clinical diagnosis in the real world often requires integrating domain knowledge, especially textual information. Conducting multimodal learning involves visual and text modalities shown as a solution, but collecting paired vision-langua...

Adaptive Vision-Guided Robotic Arm Control for Precision Pruning in Dynamic Orchard Environments

This study presents a vision-guided robotic control system for automated fruit tree pruning applications. Traditional agricultural practices rely on...

Attributes-aware Visual Emotion Representation Learning

Visual emotion analysis or recognition has gained considerable attention due to the growing interest in understanding how images can convey rich sem...

Mind the Gap: Evaluating Vision Systems in Small Data Applications

The practical application of AI tools for specific computer vision tasks relies on the "small-data regime" of hundreds to thousands of labeled sampl...

Classifying Subjective Time Perception in a Multi-robot Control Scenario Using Eye-tracking Information

As automation and mobile robotics reshape work environments, rising expectations for productivity increase cognitive demands on human operators, lea...

Diagrammatic expansion for the mutual-information rate in the realm of limited statistics

Neurons in sensory systems encode stimulus information into their stochastic spiking response. The Mutual information has been broadly applied to th...

Accuracy Enhancement in Refractive Index Sensing via Full-Spectrum Machine Learning Modeling

We present a full-spectrum machine learning framework for refractive index sensing using simulated absorption spectra from meta-grating structures c...

WoundAmbit: Bridging State-of-the-Art Semantic Segmentation and Real-World Wound Care

Chronic wounds affect a large population, particularly the elderly and diabetic patients, who often exhibit limited mobility and co-existing health ...

V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual-text processing. However, existi...

Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification

Inspired by human visual attention, deep neural networks have widely adopted attention mechanisms to learn locally discriminative attributes for cha...

Taxonomy-Aware Evaluation of Vision-Language Models

When a vision-language model (VLM) is prompted to identify an entity depicted in an image, it may answer 'I see a conifer,' rather than the specific...

A Reality Check of Vision-Language Pre-training in Radiology: Have We Progressed Using Text?

Vision-language pre-training has recently gained popularity as it allows learning rich feature representations using large-scale data sources. This ...

Security Risks in Vision-Based Beam Prediction: From Spatial Proxy Attacks to Feature Refinement

The rapid evolution towards the sixth-generation (6G) networks demands advanced beamforming techniques to address challenges in dynamic, high-mobili...

Balancing Robustness and Efficiency in Embedded DNNs Through Activation Function Selection

Machine learning-based embedded systems for safety-critical applications, such as aerospace and autonomous driving, must be robust to perturbations ...

RS-RAG: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model

Recent progress in VLMs has demonstrated impressive capabilities across a variety of tasks in the natural image domain. Motivated by these advanceme...

Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision

We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks...

SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models

Typographic attacks exploit the interplay between text and visual content in multimodal foundation models, causing misclassifications when misleadin...

Dynamic Vision Mamba

Mamba-based vision models have gained extensive attention as a result of being computationally more efficient than attention-based models. However, ...

Enhancing Compositional Reasoning in Vision-Language Models with Synthetic Preference Data

Compositionality, or correctly recognizing scenes as compositions of atomic visual concepts, remains difficult for multimodal large language models ...

Browse Categories