Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4881-4900 of 9,853 articles

FedELR: When federated learning meets learning with noisy labels.

Existing research on federated learning (FL) usually assumes that training labels are of high quality for each client, which is impractical in many real-world scenarios (e.g., noisy labels by crowd-sourced annotations), leading to dramatic performance degradation. In this work, we investigate noisy FL through the lens of early-time training phenomenon (ETP). Specifically, a key finding of this pap...

Jul 1 2025 40081270

Improving generalization of neural Vehicle Routing Problem solvers through the lens of model architecture.

Neural models produce promising results when solving Vehicle Routing Problems (VRPs), but may often fall short in generalization. Recent attempts to enhance model generalization often incur unnecessarily large training cost or cannot be directly applied to other models solving different VRP variants. To address these issues, we take a novel perspective on model architecture in this study. Specific...

Jul 1 2025 40112637
Enabling scale and rotation invariance in convolutional neural networks with retina like transformation.

Traditional convolutional neural networks (CNNs) struggle with scale and rotation transformations, resulting in reduced performance on transformed ima...

Jul 1 2025 40121784
Using large language models as decision support tools in emergency ophthalmology.

BACKGROUND: Large language models (LLMs) have shown promise in various medical applications, but their potential as decision support tools in emergenc...

Jul 1 2025 40147415
Imaging of Geographic Atrophy: A Practical Approach.

Geographic atrophy (GA) secondary to age-related macular degeneration is a chronic degenerative disease involving the retinal pigment epithelium, phot...

Jul 1 2025 40323558
CS-VLM: Compressed Sensing Attention for Efficient Vision-Language Representation Learning

Vision-Language Models (vLLMs) have emerged as powerful architectures for joint reasoning over visual and textual inputs, enabling breakthroughs in ...

GazeTarget360: Towards Gaze Target Estimation in 360-Degree for Robot Perception

Enabling robots to understand human gaze target is a crucial step to allow capabilities in downstream tasks, for example, attention estimation and m...

SurgiSR4K: A High-Resolution Endoscopic Video Dataset for Robotic-Assisted Minimally Invasive Procedures

High-resolution imaging is crucial for enhancing visual clarity and enabling precise computer-assisted guidance in minimally invasive surgery (MIS)....

Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Recent progress in multimodal reasoning has been significantly advanced by textual Chain-of-Thought (CoT), a paradigm where models conduct reasoning...

Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model

Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visu...

Visual Textualization for Image Prompted Object Detection

We propose VisTex-OVLM, a novel image prompted object detection method that introduces visual textualization -- a process that projects a few visual...

On the Domain Robustness of Contrastive Vision-Language Models

In real-world vision-language applications, practitioners increasingly rely on large, pretrained foundation models rather than custom-built solution...

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching

Driven by large-scale contrastive vision-language pre-trained models such as CLIP, recent advancements in the image-text matching task have achieved...

Sanitizing Manufacturing Dataset Labels Using Vision-Language Models

The success of machine learning models in industrial applications is heavily dependent on the quality of the datasets used to train the models. Howe...

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Vision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain. We in...

Deep Learning for Optical Misalignment Diagnostics in Multi-Lens Imaging Systems

In the rapidly evolving field of optical engineering, precise alignment of multi-lens imaging systems is critical yet challenging, as even minor mis...

Deep Learning in Mild Cognitive Impairment Diagnosis using Eye Movements and Image Content in Visual Memory Tasks

The global prevalence of dementia is projected to double by 2050, highlighting the urgent need for scalable diagnostic tools. This study utilizes di...

Revisiting CroPA: A Reproducibility Study and Enhancements for Cross-Prompt Adversarial Transferability in Vision-Language Models

Large Vision-Language Models (VLMs) have revolutionized computer vision, enabling tasks such as image classification, captioning, and visual questio...

MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering

Medical visual question answering (MedVQA) plays a vital role in clinical decision-making by providing contextually rich answers to image-based quer...

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

This work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule...

Browse Categories