Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5421-5440 of 9,853 articles

SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting

Open-world object counting leverages the robust text-image alignment of pre-trained vision-language models (VLMs) to enable counting of arbitrary categories in images specified by textual queries. However, widely adopted naive fine-tuning strategies concentrate exclusively on text-image consistency for categories contained in training, which leads to limited generalizability for unseen categorie...

Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation

We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environment or situation from an alternative viewpoint, is a key benchmark for human-level visual understanding, essential for environmental interaction and collaboration with autonomous agents. Despite advancements in spatial r...

DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs

We present DyMU, an efficient, training-free framework that dynamically reduces the computational burden of vision-language models (VLMs) while main...

Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism

Image description generation is essential for accessibility and AI understanding of visual content. Recent advancements in deep learning have signif...

Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes

Streetscapes are an essential component of urban space. Their assessment is presently either limited to morphometric properties of their mass skelet...

FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing

In recent years, large-scale vision-language models (VLMs) like CLIP have gained attention for their zero-shot inference using instructional text pr...

Deep Multi-modal Breast Cancer Detection Network

Automated breast cancer detection via computer vision techniques is challenging due to the complex nature of breast tissue, the subtle appearance of...

A Clinician-Friendly Platform for Ophthalmic Image Analysis Without Technical Barriers

Artificial intelligence (AI) shows remarkable potential in medical imaging diagnostics, yet most current models require retraining when applied acro...

Integrating Non-Linear Radon Transformation for Diabetic Retinopathy Grading

Diabetic retinopathy is a serious ocular complication that poses a significant threat to patients' vision and overall health. Early detection and ac...

Observability conditions for neural state-space models with eigenvalues and their roots of unity

We operate through the lens of ordinary differential equations and control theory to study the concept of observability in the context of neural sta...

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

Preference alignment through Direct Preference Optimization (DPO) has demonstrated significant effectiveness in aligning multimodal large language m...

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Recognizing and reasoning about occluded (partially or fully hidden) objects is vital to understanding visual scenes, as occlusions frequently occur...

Manifold Induced Biases for Zero-shot and Few-shot Detection of Generated Images

Distinguishing between real and AI-generated images, commonly referred to as 'image detection', presents a timely and significant challenge. Despite...

Shifts in Doctors' Eye Movements Between Real and AI-Generated Medical Images

Eye-tracking analysis plays a vital role in medical imaging, providing key insights into how radiologists visually interpret and diagnose clinical c...

ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages

Vision Transformers (ViTs) have revolutionized computer vision by leveraging self-attention to model long-range dependencies. However, ViTs face cha...

Real-Time Sleepiness Detection for Driver State Monitoring System

A driver face monitoring system can detect driver fatigue, which is a significant factor in many accidents, using computer vision techniques. In thi...

OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding

The practical deployment of medical vision-language models (Med-VLMs) necessitates seamless integration of textual data with diverse visual modaliti...

NVSMask3D: Hard Visual Prompting with Camera Pose Interpolation for 3D Open Vocabulary Instance Segmentation

Vision-language models (VLMs) have demonstrated impressive zero-shot transfer capabilities in image-level visual perception tasks. However, they fal...

Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding

Vision Large Language Models (VLLMs) have demonstrated impressive capabilities in general visual tasks such as image captioning and visual question ...

LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation

Zero-shot referring image segmentation aims to locate and segment the target region based on a referring expression, with the primary challenge of a...

Browse Categories