Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5241-5260 of 9,853 articles

PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning

Retinal foundation models have significantly advanced retinal image analysis by leveraging self-supervised learning to reduce dependence on labeled data while achieving strong generalization. Many recent approaches enhance retinal image understanding using report supervision, but obtaining clinical reports is often costly and challenging. In contrast, metadata (e.g., age, gender) is widely avail...

Bayesian Deep Learning Approaches for Uncertainty-Aware Retinal OCT Image Segmentation for Multiple Sclerosis

Optical Coherence Tomography (OCT) provides valuable insights in ophthalmology, cardiology, and neurology due to high-resolution, cross-sectional images of the retina. One critical task for ophthalmologists using OCT is delineation of retinal layers within scans. This process is time-consuming and prone to human bias, affecting the accuracy and reliability of diagnoses. Previous efforts to autom...

IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests

Although large Vision-Language Models (VLMs) have demonstrated remarkable performance in a wide range of multimodal tasks, their true reasoning capa...

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation

Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence-par...

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding

Visual grounding is essential for precise perception and reasoning in multimodal large language models (MLLMs), especially in medical imaging domain...

ElderFallGuard: Real-Time IoT and Computer Vision-Based Fall Detection System for Elderly Safety

For the elderly population, falls pose a serious and increasing risk of serious injury and loss of independence. In order to overcome this difficult...

UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings

Current research has explored vision-language models for multi-modal embedding tasks, such as information retrieval, visual grounding, and classific...

MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark

Document retrieval is an important task for search and Retrieval-Augmented Generation (RAG) applications. Large Language Models (LLMs) have contribu...

Visual Planning: Let's Think Only with Images

Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across div...

Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild

To perform autonomous visual search for environmental monitoring, a robot may leverage satellite imagery as a prior map. This can help inform coarse...

Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models

Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image captioning and video question answering. However, real-t...

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions

Vision-Language-Action (VLA) models have recently become highly prominent in the field of robotics. Leveraging vision-language foundation models tra...

PhiNet v2: A Mask-Free Brain-Inspired Vision Foundation Model from Video

Recent advances in self-supervised learning (SSL) have revolutionized computer vision through innovative architectures and learning objectives, yet ...

What's Inside Your Diffusion Model? A Score-Based Riemannian Metric to Explore the Data Manifold

Recent advances in diffusion models have demonstrated their remarkable ability to capture complex image distributions, but the geometric properties ...

Unifying Segment Anything in Microscopy with Multimodal Large Language Model

Accurate segmentation of regions of interest in biomedical images holds substantial value in image analysis. Although several foundation models for ...

Benchmarking performance, explainability, and evaluation strategies of vision-language models for surgery: Challenges and opportunities

Minimally invasive surgery (MIS) presents significant visual and technical challenges, including surgical instrument classification and understandin...

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are cap...

End-to-End Vision Tokenizer Tuning

Existing vision tokenization isolates the optimization of vision tokenizers from downstream training, implicitly assuming the visual tokens can gene...

AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection

Universal visual anomaly detection aims to identify anomalies from novel or unseen vision domains without additional fine-tuning, which is critical ...

Comparison of 11 intraocular lens power calculation formulas in eyes undergoing simultaneous cataract surgery and Descemet membrane endothelial keratoplasty.

PURPOSE: To evaluate the accuracy of 11 intraocular lens (IOL) calculation formulas in eyes undergoing Descemet membrane endothelial keratoplasty (DME...

May 15 2025 40381708
Browse Categories