Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4481-4500 of 9,853 articles

MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image Retrieval

Composed Image Retrieval (CIR) aims to retrieve target images based on a reference image and modified texts. However, existing methods often struggle to extract the correct semantic cues from the reference image that best reflect the user's intent under textual modification prompts, resulting in interference from irrelevant visual noise. In this paper, we propose a novel Multi-level Vision Selecti...

Mar 18 2026 2603.17360v1

Harnessing the Power of Foundation Models for Accurate Material Classification

Material classification has emerged as a critical task in computer vision and graphics, supporting the assignment of accurate material properties to a wide range of digital and real-world applications. While traditionally framed as an image classification task, this domain faces significant challenges due to the scarcity of annotated data, limiting the accuracy and generalizability of trained mode...

Mar 18 2026 2603.17390v1
EI: Early Intervention for Multimodal Imaging based Disease Recognition

Current methods for multimodal medical imaging based disease recognition face two major challenges. First, the prevailing "fusion after unimodal image...

Mar 18 2026 2603.17514v1
Interpretable Cross-Domain Few-Shot Learning with Rectified Target-Domain Local Alignment

Cross-Domain Few-Shot Learning (CDFSL) adapts models trained with large-scale general data (source domain) to downstream target domains with only scar...

Mar 18 2026 2603.17655v1
Eye image segmentation using visual and concept prompts with Segment Anything Model 3 (SAM3)

Previous work has reported that vision foundation models show promising zero-shot performance in eye image segmentation. Here we examine whether the l...

Mar 18 2026 2603.17715v1
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models

Vision-Language Models (VLMs) are increasingly applied to robotic perception and manipulation, yet their ability to infer physical properties required...

Mar 17 2026 2603.16958v1
LLM-Powered Flood Depth Estimation from Social Media Imagery: A Vision-Language Model Framework with Mechanistic Interpretability for Transportation Resilience

Urban flooding poses an escalating threat to transportation network continuity, yet no operational system currently provides real-time, street-level f...

Mar 17 2026 2603.17108v1
A Lensless Polarization Camera

Polarization imaging is a technique that creates a pixel map of the polarization state in a scene. Although invisible to the human eye, polarization c...

Mar 17 2026 2603.17156v1
BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images

We present BEV-SLD, a LiDAR global localization method building on the Scene Landmark Detection (SLD) concept. Unlike scene-agnostic pipelines, our se...

Mar 17 2026 2603.17159v1
Visual Product Search Benchmark

Reliable product identification from images is a critical requirement in industrial and commercial applications, particularly in maintenance, procurem...

Mar 17 2026 2603.17186v1
EchoAtlas: A Conversational, Multi-View Vision-Language Foundation Model for Echocardiography Interpretation and Clinical Reasoning

Echocardiography is the most widely used cardiac imaging modality, yet artificial intelligence-enabled interpretation remains limited by the inability...

PathGLS: Evaluating Pathology Vision-Language Models without Ground Truth through Multi-Dimensional Consistency

Vision-Language Models (VLMs) offer significant potential in computational pathology by enabling interpretable image analysis, automated reporting, an...

Mar 17 2026 2603.16113v1
AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models

Artificial Intelligence (AI) workloads drive a rapid expansion of high-performance computing (HPC) infrastructures and increase their power and energy...

Mar 17 2026 2603.16164v1
KidsNanny: A Two-Stage Multimodal Content Moderation Pipeline Integrating Visual Classification, Object Detection, OCR, and Contextual Reasoning for Child Safety

We present KidsNanny, a two-stage multimodal content moderation architecture for child safety. Stage 1 combines a vision transformer (ViT) with an obj...

Mar 17 2026 2603.16181v1
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models

Vision-language process reward models (VL-PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test-time scal...

Mar 17 2026 2603.16253v1
SF-Mamba: Rethinking State Space Model for Vision

The realm of Mamba for vision has been advanced in recent years to strike for the alternatives of Vision Transformers (ViTs) that suffer from the quad...

Mar 17 2026 2603.16423v1
On the Transfer of Collinearity to Computer Vision

Collinearity is a visual perception phenomenon in the human brain that amplifies spatially aligned edges arranged along a straight line. However, it i...

Mar 17 2026 2603.16592v1
HeBA: Heterogeneous Bottleneck Adapters for Robust Vision-Language Models

Adapting large-scale Vision-Language Models (VLMs) like CLIP to downstream tasks often suffers from a "one-size-fits-all" architectural approach, wher...

Mar 17 2026 2603.16653v1
vAccSOL: Efficient and Transparent AI Vision Offloading for Mobile Robots

Mobile robots are increasingly deployed for inspection, patrol, and search-and-rescue operations, relying on computer vision for perception, navigatio...

Mar 17 2026 2603.16685v1
Multimodal Machine Learning for Glaucoma Detection in a Sub-Saharan African Clinical Population

Purpose: To evaluate the performance of machine learning models for automated glaucoma detection using multimodal clinical, structural, and functional...

Browse Categories