Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4061-4080 of 9,853 articles

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it remains unclear whether their success reflects genuine causal understanding or sophisticated pattern matching over visual-textual correlations. Inspired by Russell's inductivist turkey, we introduce Counterfactual-World (CF-World), a counterfactual be...

Jun 23 2026 2606.24548v2

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified architecture enables strong performance on tasks like image captioning, visual question answering, and multimodal dialogue, often in zero- and few-shot settings. Their general-purpose capabilities and flexible interfaces make MLLMs a promising fou...

Jun 23 2026 2606.25084v1
DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model

Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follo...

Jun 23 2026 2606.24051v1
TuringViT: Making SOTA Vision Transformers Accessible to All

Modern VLMs and VLA systems commonly adopt off-the-shelf ViTs such as SigLIP2 as visual encoders, but diverse downstream requirements in latency, temp...

Jun 23 2026 2606.24253v1
Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it r...

Jun 23 2026 2606.24548v1
HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models

Vision-Language Models (VLMs) are typically pre-trained on large-scale image-text datasets to capture semantic correspondences between visual content ...

Jun 22 2026 2606.23843v1
E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis

While Vision-Language Models (VLMs) show great promise in volumetric medical report generation, they frequently suffer from visual hallucinations and ...

Jun 22 2026 2606.23888v1
How does AI detect diabetic retinopathy from retinal photos? A heatmap analysis of 54 deep learning models

Purpose: To investigate how artificial intelligence (AI) systems detect referrable diabetic retinopathy (DR) from retinal photographs by analysing hea...

Sequential Deep Learning to Predict Non-Central to Central Geographic Atrophy Progression from OCT Imaging

Purpose: To develop and validate a temporal deep learning framework for predicting geographic atrophy (GA) progression across multi-year horizons usin...

DE-FIVE: Detecting Malicious Image Prompts via Fourier Features and Image Vector Embeddings

Vision language models (VLMs) employ both visual and textual modalities to enable advanced vision-language inference. However, incorporating visual mo...

Jun 22 2026 2606.22779v1
T-VSS: Test-Time Visual Subspace Steering for Adversarial Robustness of Vision-Language Models

Vision-language models (VLMs) achieve strong zero-shot recognition, but they remain highly vulnerable to adversarial perturbations. Recent test-time a...

Jun 22 2026 2606.23132v1
Sublinearly Structured Deep Neural Networks Achieve Feature Learning Consistency for Compositional Functions

Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations ...

Jun 22 2026 2606.23477v1
Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies

Automated diagnosis of 3D brain CT scans is essential for critical care, yet it remains challenging due to the heavy reliance on manual annotations an...

Jun 22 2026 2606.23494v1
Benchmarking Vision-Language Models for Microscopic Plant Image Understanding

Microscopic imaging provides essential visual evidence for studying plant biology and pathology at the cellular and subcellular levels. However, exist...

Jun 21 2026 2606.22497v1
MAPS: Multi-Anchor Projection Similarity for Joint Vision-Language Geo-Localization

Humans localize places by integrating perceptual cues from vision with semantic reasoning from language, forming a scene understanding that is both in...

Jun 21 2026 2606.22543v1
Extraction of Glaucoma Diagnosis, Type, and Severity from Clinical Notes using Secure Cloud-based Large Language Models

Purpose: To evaluate the performance of secure cloud-based large language models (LLMs) in extracting glaucoma diagnosis, type, and severity from free...

Benchmarking attention-based methods for vision transformers' interpretability in retinal fundus imaging

Deep learning models based on Vision Transformers (ViTs) have shown strong performance in retinal fundus imaging, but their interpretability remains p...

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding

Recently, vision-language models (VLMs) have made significant progress in 3D scene understanding, driving advances in applications such as embodied in...

Jun 18 2026 2606.19776v1
CSWinUNETR: Segmentation of Thin Anatomical Structures in Medical Images

Accurate segmentation of thin, tortuous anatomical structures, such as retinal vessels, cerebral vasculature, and facial wrinkles, remains challenging...

Jun 18 2026 2606.19824v1
Vision-Reasoning-Guided Occlusion Removal from Light Fields

Occlusion-robust scene recovery remains a major challenge in computational imaging, particularly in natural environments where dense foreground vegeta...

Jun 18 2026 2606.19985v1
Browse Categories