Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4321-4340 of 9,853 articles

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?

Recent advancements in self-supervised learning have led to powerful surgical vision encoders capable of spatiotemporal understanding. However, extending these visual foundations to multi-modal reasoning tasks is severely bottlenecked by the prohibitive cost of expert textual annotations. To overcome this scalability limitation, we introduce \textbf{LIME}, a large-scale multi-modal dataset derived...

Apr 20 2026 2604.18134v1

Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning

Accurate prognostication and risk estimation are essential for guiding clinical decision-making and optimizing patient management. While radiologist-assessed features from CT scans provide valuable indicators of disease severity and outcomes, interpreting such images requires expert knowledge, and translating rich visual information into textual summaries inevitably leads to information loss. In t...

Apr 20 2026 2604.18250v1
Long-Text-to-Image Generation via Compositional Prompt Decomposition

While modern text-to-image (T2I) models excel at generating images from intricate prompts, they struggle to capture the key details when the inputs ar...

Apr 20 2026 2604.18258v1
Geometry-Guided 3D Visual Token Pruning for Video-Language Models

Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent ...

Apr 20 2026 2604.18260v1
Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models

Change visual question answering (Change VQA) addresses the problem of answering natural-language questions about semantic changes between bi-temporal...

Apr 20 2026 2604.18429v1
S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models

Vision-Language Models (VLMs) have demonstrated remarkable progress in single-image understanding, yet effective reasoning across multiple images rema...

Apr 20 2026 2604.18512v1
Advancing Vision Transformer with Enhanced Spatial Priors

In recent years, the Vision Transformer (ViT) has garnered significant attention within the computer vision community. However, the core component of ...

Apr 20 2026 2604.18549v1
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability

Despite recent progress, vision-language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, w...

Apr 20 2026 2604.18573v1
DREAM: Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion for Expert Precision Medical Report Generation

Automating medical reports for retinal images requires a sophisticated blend of visual pattern recognition and deep clinical knowledge. Current Large ...

Apr 19 2026 2604.17209v1
Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability

Vision-Language Models (VLMs) achieve strong cross-modal performance, yet recent evidence suggests they over-rely on textual descriptions while under-...

Apr 19 2026 2604.17217v1
Robust Diabetic Retinopathy Grading Using Dual-Resolution Attention-Based Deep Learning with Ordinal Regression

Diabetic retinopathy (DR) is a leading cause of vision impairment worldwide, and automated grading systems play a crucial role in large-scale screenin...

Apr 19 2026 2604.17341v1
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models

Recent advances in Vision-Language Models (VLMs) have substantially enhanced their ability across multimodal video understanding benchmarks spanning t...

Apr 19 2026 2604.17375v1
RS-HyRe-R1: A Hybrid Reward Mechanism to Overcome Perceptual Inertia for Remote Sensing Images Understanding

Reinforcement learning (RL) post-training substantially improves remote sensing vision-language models (RS-VLMs). However, when handling complex remot...

Apr 19 2026 2604.17504v1
PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation

Peripheral Blood Smear (PBS) is a critical microscopic examination in hematopathology that yields whole-slide imaging (WSI). Unlike solid tissue patho...

Apr 19 2026 2604.17570v1
Imbalance-Aware Optimal Transport Learning for Cost-Effective Diabetic Retinopathy Screening

Abstract Background Diabetic Retinopathy (DR) is one of the leading cause of vision loss and blindness. AI models have been instrumental in providing ...

CPU Optimization of a Monocular 3D Biomechanics Pipeline for Low-Resource Deployment

Markerless 3D movement analysis from monocular video enables accessible biomechanical assessment in clinical and sports settings. However, most resear...

Apr 17 2026 2604.15665v1
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow

Vision-Language Models (VLMs) have demonstrated strong capability in a wide range of tasks such as visual recognition, document parsing, and visual gr...

Apr 17 2026 2604.15809v1
Learning to Look before Learning to Like: Incorporating Human Visual Cognition into Aesthetic Quality Assessment

Automated Aesthetic Quality Assessment (AQA) treats images primarily as static pixel vectors, aligning predictions with human-rating scores largely th...

Apr 17 2026 2604.15853v1
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap

Reasoning in vision-language models (VLMs) has recently attracted significant attention due to its broad applicability across diverse downstream tasks...

Apr 17 2026 2604.16256v1
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

Retrieval-Augmented Generation (RAG) extends Large Vision-Language Models (LVLMs) with external visual knowledge. However, existing visual RAG systems...

Apr 16 2026 2604.14967v2
Browse Categories