Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3781-3800 of 9,853 articles

ChartProbe: A Diagnostic Study on Visual Reasoning through Perception, Grounding, and Simple Reasoning

Vision-language models (VLMs) remain unreliable on chart questions that require reasoning over visual quantities, and this weakness is usually attributed to a reasoning deficit and addressed with more reasoning supervision. We ask whether the difficulty lies in reasoning itself, or in the simpler skills that reasoning operates on: reading the plotted elements (\emph{perception}), locating them and...

Aug 13 2026 2608.13766v1

The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature extraction difficult. In this study, a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework is ...

Aug 13 2026 2608.12677v1
MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty ...

Aug 13 2026 2608.13463v1
A Vision-Language Framework for Predicting Brain Tumor Recurrence from Multimodal, Longitudinal Patient Data

Accurate prediction of tumor recurrence in brain tumor patients following surgery is essential for optimizing adjuvant therapy, response assessment, a...

Vision and Language Models for Classifying Maxillary Sinus Disease on Cone-Beam Computed Tomography: A Transparent Multimodal Benchmark

Background: Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abn...

Can Vision Models Read the Radar Display? On the Feasibility of Radar Imagery for Air Traffic Complexity Estimation

Air traffic controllers perceive traffic complexity through the radar display, suggesting that a computer vision model operating on the same imagery m...

Aug 12 2026 2608.11810v1
Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation

In vision-based behavior cloning (BC), conventional image augmentations such as Random Crop and Color Jitter often fall short under substantial visual...

Aug 12 2026 2608.11870v1
CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data

Omics datasets, particularly single-cell RNA sequencing data, are high-dimensional, sparse, noisy, and dominated by zero values, making faithful low-d...

Aug 11 2026 2608.11269v1
Automatic Field-of-View Adjustment for a View-Expansive Microscope via LSTM-Based Gaze and Pipette Motion Interpretation

Intracytoplasmic sperm injection (ICSI) operators frequently adjust the field-of-view (FOV) during procedures, which interrupts workflow and increases...

Aug 11 2026 2608.10401v1
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their...

Aug 11 2026 2608.10513v1
Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training

While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent ...

Aug 11 2026 2608.10522v1
MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain chall...

Aug 11 2026 2608.10635v1
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri...

Aug 11 2026 2608.10636v1
MMArt A Multi-Perspective Multimodal Dataset for Visual Art Understanding

Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface...

Aug 11 2026 2608.10706v1
Where To Look? : Causal Tracing of Vision Encoders in VLM

Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information act...

Aug 11 2026 2608.10758v1
Shaping the notion of #wellbeing in the therapy culture context: an analysis through Instagram narratives

The aim of this paper is to (1) identify textual and visual themes and sub-themes associated with the #wellbeing hashtag on Instagram, (2) assess thei...

Aug 11 2026 2608.10793v1
Modelling Geographic Atrophy Progression using Implicit Neural Representations

Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irreversible atrop...

Aug 11 2026 2608.10807v1
Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models

Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer....

Aug 11 2026 2608.10824v1
Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragil...

Aug 11 2026 2608.10864v1
When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

Attribute hallucination---where vision-language models (VLMs) correctly identify an object but mischaracterize its properties---is prevalent yet mecha...

Aug 11 2026 2608.11024v1
Browse Categories