Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4941-4960 of 9,853 articles

Boundary-Aware Vision Transformer for Angiography Vascular Network Segmentation

Accurate segmentation of vascular structures in coronary angiography remains a core challenge in medical image analysis due to the complexity of elongated, thin, and low-contrast vessels. Classical convolutional neural networks (CNNs) often fail to preserve topological continuity, while recent Vision Transformer (ViT)-based models, although strong in global context modeling, lack precise boundar...

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

Vision-Language Models (VLMs) face significant challenges when dealing with the diverse resolutions and aspect ratios of real-world images, as most existing models rely on fixed, low-resolution inputs. While recent studies have explored integrating native resolution visual encoding to improve model performance, such efforts remain fragmented and lack a systematic framework within the open-source...

NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models

Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable capabilities in understanding relationships between visual and textual data ...

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation

Large vision-language models (LVLMs) have shown remarkable capabilities across a wide range of multimodal tasks. However, they remain prone to visua...

Inference-Time Gaze Refinement for Micro-Expression Recognition: Enhancing Event-Based Eye Tracking with Motion-Aware Post-Processing

Event-based eye tracking holds significant promise for fine-grained cognitive state inference, offering high temporal resolution and robustness to m...

Comparative Analysis of Deep Learning Strategies for Hypertensive Retinopathy Detection from Fundus Images: From Scratch and Pre-trained Models

This paper presents a comparative analysis of deep learning strategies for detecting hypertensive retinopathy from fundus images, a central task in ...

Feeling Machines: Ethics, Culture, and the Rise of Emotional AI

This paper explores the growing presence of emotionally responsive artificial intelligence through a critical and interdisciplinary lens. Bringing t...

Optimized Spectral Fault Receptive Fields for Diagnosis-Informed Prognosis

This paper introduces Spectral Fault Receptive Fields (SFRFs), a biologically inspired technique for degradation state assessment in bearing fault d...

VGR: Visual Grounded Reasoning

In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inher...

VGR: Visual Grounded Reasoning

In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inher...

Evaluating Sensitivity Parameters in Smartphone-Based Gaze Estimation: A Comparative Study of Appearance-Based and Infrared Eye Trackers

This study evaluates a smartphone-based, deep-learning eye-tracking algorithm by comparing its performance against a commercial infrared-based eye t...

Evaluating Sensitivity Parameters in Smartphone-Based Gaze Estimation: A Comparative Study of Appearance-Based and Infrared Eye Trackers

This study evaluates a smartphone-based, deep-learning eye-tracking algorithm by comparing its performance against a commercial infrared-based eye t...

Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation

Vision-Language Translation (VLT) is a challenging task that requires accurately recognizing multilingual text embedded in images and translating it...

MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in interpreting visual layouts and text. However, a significant challenge re...

EasyARC: Evaluating Vision Language Models on True Visual Reasoning

Building on recent advances in language-based reasoning models, we explore multimodal reasoning that integrates vision and text. Existing multimodal...

EyeSim-VQA: A Free-Energy-Guided Eye Simulation Framework for Video Quality Assessment

Free-energy-guided self-repair mechanisms have shown promising results in image quality assessment (IQA), but remain under-explored in video quality...

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

In multimodal large language models (MLLMs), the length of input visual tokens is often significantly greater than that of their textual counterpart...

In-Hand Object Pose Estimation via Visual-Tactile Fusion

Accurate in-hand pose estimation is crucial for robotic object manipulation, but visual occlusion remains a major challenge for vision-based approac...

Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

Vision-and-language navigation (VLN) is a long-standing challenge in autonomous robotics, aiming to empower agents with the ability to follow human ...

PiPViT: Patch-based Visual Interpretable Prototypes for Retinal Image Analysis

Background and Objective: Prototype-based methods improve interpretability by learning fine-grained part-prototypes; however, their visualization in...

Browse Categories