Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4161-4180 of 9,853 articles

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness against adversarial attacks, particularly those exploiting both modalities, remains underexplored, posing risks to critical applications like autonomous driving and content moderat...

May 26 2026 2605.26501v1

Re-M3Dr: Rebalanced MultiModal Mean Deviation Regression

Mean Deviation (MD) is a critical metric for assessing visual field loss in ophthalmology. While previous work has focused solely on predicting MD from Optical Coherence Tomography (OCT), it is intuitive to assume that combining OCT with another imaging of fundus photography (FP) could improve performance, as two ophthalmic medical imaging provide complementary information. This is particularly ex...

May 26 2026 2605.26513v1
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention ...

May 26 2026 2605.26636v1
DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding

Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to ...

May 26 2026 2605.26656v1
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning

Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they often reason in language and lose the fine-grained geometry n...

May 26 2026 2605.27310v1
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection

Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularl...

May 26 2026 2605.27348v1
Benchmarking Convolutional, Transformer, Hybrid, and Vision Language Models for Multi Disease Retinal Screening

Modern deep learning offers powerful tools for automated retinal screening, but it remains unclear how different visual model families compare in real...

May 25 2026 2605.26283v1
Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning

Parameter-efficient adaptation of vision-language foundation models is crucial for precise multimodal understanding of biomedical images, yet existing...

May 25 2026 2605.26292v1
NightSight: Passive Computation for Navigation in Dark Using Events

Small aerial robots are particularly well-suited for search and rescue in confined and hazardous environments due to their agility, low cost, and abil...

May 25 2026 2605.26330v1
Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence

Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of geometric knowledge: holistic 3D structural per...

May 25 2026 2605.25334v1
Toward Native Multimodal Modeling: A Roadmap

Multimodal modeling represents a vital step from modality-agnostic reasoning toward world modeling. While early approaches predominantly rely on late-...

May 25 2026 2605.25343v1
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models

Adapting large vision-language models (VLMs) such as CLIP to downstream tasks remains challenging, as full fine-tuning is computationally prohibitive ...

May 25 2026 2605.25479v1
CMAP: Cross-Modal Adaptive Prompting for Multi-Domain Task-Incremental Learning

Multi-domain task-incremental learning requires a model to sequentially acquire knowledge across visually diverse domains without forgetting prior tas...

May 25 2026 2605.25708v1
[CLS] is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive Aggregation

Vision-Language Models such as CLIP exhibit strong zero-shot recognition capability by aligning images with textual concepts, yet they often underperf...

May 25 2026 2605.25821v1
AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language des...

May 25 2026 2605.25901v1
RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing

Urgent suspected colorectal cancer (CRC) referrals create operational bottlenecks because semi-structured clinical documents often require manual revi...

May 25 2026 2605.25956v1
Context-driven Missing-Modality Learning for Robust Medical Diagnosis with Image-Tabular Data

While multimodal data integrating diverse imaging and clinical tabular records is crucial for accurate medical diagnosis, the arbitrary absence of spe...

May 25 2026 2605.25968v1
DeltaCam: Differential Intrinsic Camera Modeling for Video Generation

Incorporating camera intrinsics into video generation models offers a principled way to control not only scene dynamics but also the imaging process t...

May 24 2026 2605.25266v1
Design and Validation of an AI-Assisted Sequential Screening Framework for Psychological Distress in Glaucoma

Purpose: Psychological distress is highly prevalent in glaucoma and is associated with worse adherence, reduced quality of life, and faster disease pr...

A Competitive Framework for Modeling EEG Microstate Durations

Background. This study examines a competition based model (Cmodel) designed to capture the temporal dynamics of successive brain microstates derived f...

Browse Categories