Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5721-5740 of 9,853 articles

VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models

Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Recently, reward models (RMs) have become increasingly pivotal in the reasoning process. Specifically, process RMs evaluate each reasoning step, outcome RMs focus on the assessment of reasoning results, and critique RMs ...

A Novel Ophthalmic Benchmark for Evaluating Multimodal Large Language Models with Fundus Photographs and OCT Images

In recent years, large language models (LLMs) have demonstrated remarkable potential across various medical applications. Building on this foundation, multimodal large language models (MLLMs) integrate LLMs with visual models to process diverse inputs, including clinical data and medical images. In ophthalmology, LLMs have been explored for analyzing optical coherence tomography (OCT) reports, a...

Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning

Multimodal large language models (MLLMs), built on large-scale pre-trained vision towers and language models, have shown great capabilities in multi...

Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment

Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations b...

Eye-in-Finger: Smart Fingers for Delicate Assembly and Disassembly of LEGO

Manipulation and insertion of small and tight-toleranced objects in robotic assembly remain a critical challenge for vision-based robotics systems d...

HierDAMap: Towards Universal Domain Adaptive BEV Mapping via Hierarchical Perspective Priors

The exploration of Bird's-Eye View (BEV) mapping technology has driven significant innovation in visual perception technology for autonomous driving...

X-GAN: A Generative AI-Powered Unsupervised Model for High-Precision Segmentation of Retinal Main Vessels toward Early Detection of Glaucoma

Structural changes in main retinal blood vessels serve as critical biomarkers for the onset and progression of glaucoma. Identifying these vessels i...

Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On

Interpretability in Vision-Language Models (VLMs) is crucial for trust, debugging, and decision-making in high-stakes applications. We introduce Pix...

Dynamic Updates for Language Adaptation in Visual-Language Tracking

The consistency between the semantic information provided by the multi-modal reference and the tracked object is crucial for visual-language (VL) tr...

CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model

Chinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fai...

Spectral State Space Model for Rotation-Invariant Visual Representation Learning

State Space Models (SSMs) have recently emerged as an alternative to Vision Transformers (ViTs) due to their unique ability of modeling global relat...

GeoLangBind: Unifying Earth Observation with Agglomerative Vision-Language Foundation Models

Earth observation (EO) data, collected from diverse sensors with varying imaging principles, present significant challenges in creating unified anal...

Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

Visual grounding seeks to localize the image region corresponding to a free-form text description. Recently, the strong multimodal capabilities of L...

A 7T fMRI dataset of synthetic images for out-of-distribution modeling of vision

Large-scale visual neural datasets such as the Natural Scenes Dataset (NSD) are boosting NeuroAI research by enabling computational models of the br...

VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide lim...

Treble Counterfactual VLMs: A Causal Approach to Hallucination

Vision-Language Models (VLMs) have advanced multi-modal tasks like image captioning, visual question answering, and reasoning. However, they often g...

Towards Universal Text-driven CT Image Segmentation

Computed tomography (CT) is extensively used for accurate visualization and segmentation of organs and lesions. While deep learning models such as c...

Integrating Frequency-Domain Representations with Low-Rank Adaptation in Vision-Language Models

Situational awareness applications rely heavily on real-time processing of visual and textual data to provide actionable insights. Vision language m...

GrInAdapt: Scaling Retinal Vessel Structural Map Segmentation Through Grounding, Integrating and Adapting Multi-device, Multi-site, and Multi-modal Fundus Domains

Retinal vessel segmentation is critical for diagnosing ocular conditions, yet current deep learning methods are limited by modality-specific challen...

Towards Understanding the Use of MLLM-Enabled Applications for Visual Interpretation by Blind and Low Vision People

Blind and Low Vision (BLV) people have adopted AI-powered visual interpretation applications to address their daily needs. While these applications ...

Browse Categories