Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5081-5100 of 9,853 articles

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT

Understanding the physical world - governed by laws of motion, spatial relations, and causality - poses a fundamental challenge for multimodal large language models (MLLMs). While recent advances such as OpenAI o3 and GPT-4o demonstrate impressive perceptual and reasoning capabilities, our investigation reveals these models struggle profoundly with visual physical reasoning, failing to grasp bas...

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?

Vision-language models (VLMs) exhibit strong zero-shot generalization on natural images and show early promise in interpretable medical image analysis. However, existing benchmarks do not systematically evaluate whether these models truly reason like human clinicians or merely imitate superficial patterns. To address this gap, we propose DrVD-Bench, the first multimodal benchmark for clinical vi...

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation

The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches...

DeepDTAGen: a multitask deep learning framework for drug-target affinity prediction and target-aware drugs generation.

Identifying novel drugs that can interact with target proteins is a highly challenging, time-consuming, and costly task in drug discovery and developm...

May 30 2025 40447614
MCOA: A Comprehensive Multimodal Dataset for Advancing Deep Learning in Corneal Opacity Assessment.

Corneal opacity remains a major global cause of vision impairment. Its severity is typically assessed subjectively by clinicians using slit lamp exami...

May 30 2025 40447652
DGIQA: Depth-guided Feature Attention and Refinement for Generalizable Image Quality Assessment

A long-held challenge in no-reference image quality assessment (NR-IQA) learning from human subjective perception is the lack of objective generaliz...

Improved Accuracy in Pelvic Tumor Resections Using a Real-Time Vision-Guided Surgical System

Pelvic bone tumor resections remain significantly challenging due to complex three-dimensional anatomy and limited surgical visualization. Current n...

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks, yet they often strug...

Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint

Rebus puzzles, visual riddles that encode language through imagery, spatial arrangement, and symbolic substitution, pose a unique challenge to curre...

CLDTracker: A Comprehensive Language Description for Visual Tracking

VOT remains a fundamental yet challenging task in computer vision due to dynamic appearance changes, occlusions, and background clutter. Traditional...

DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers

Visual Prompt Tuning (VPT) has become a promising solution for Parameter-Efficient Fine-Tuning (PEFT) approach for Vision Transformer (ViT) models b...

DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers

Visual Prompt Tuning (VPT) has become a promising solution for Parameter-Efficient Fine-Tuning (PEFT) approach for Vision Transformer (ViT) models b...

Position Paper: Metadata Enrichment Model: Integrating Neural Networks and Semantic Knowledge Graphs for Cultural Heritage Applications

The digitization of cultural heritage collections has opened new directions for research, yet the lack of enriched metadata poses a substantial chal...

Vision-Integrated High-Quality Neural Speech Coding

This paper proposes a novel vision-integrated neural speech codec (VNSC), which aims to enhance speech coding quality by leveraging visual modality ...

VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?

Recent studies have shown that long chain-of-thought (CoT) reasoning can significantly enhance the performance of large language models (LLMs) on co...

QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining

Multimodal Large Language Models (MLLMs) encode images into visual tokens, aligning visual and textual signals within a shared latent space to facil...

Vision-Based Assistive Technologies for People with Cerebral Visual Impairment: A Review and Focus Study

Over the past decade, considerable research has investigated Vision-Based Assistive Technologies (VBAT) to support people with vision impairments to...

Cultural Evaluations of Vision-Language Models Have a Lot to Learn from Cultural Theory

Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VL...

MIAS-SAM: Medical Image Anomaly Segmentation without thresholding

This paper presents MIAS-SAM, a novel approach for the segmentation of anomalous regions in medical images. MIAS-SAM uses a patch-based memory bank ...

Adversarially Robust AI-Generated Image Detection for Free: An Information Theoretic Perspective

Rapid advances in Artificial Intelligence Generated Images (AIGI) have facilitated malicious use, such as forgery and misinformation. Therefore, num...

Browse Categories