Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3941-3960 of 9,853 articles

SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning

Vision Language Models (VLMs) demonstrate strong perceptual abilities but remain limited in tasks requiring analytical reasoning across multiple visual states, such as multi-image comparison, change detection, and multi-step visual inference. These capabilities are critical for real-world multimodal applications where reasoning must be grounded in systematic differences between visual contexts. Ho...

Jul 15 2026 2607.14333v1

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support visual and semantic generalization. Co-training on web image-text data, a common remedy, does not prevent this; it applies language ...

Jul 15 2026 2607.13429v1
Video to All-in-focus Image Reconstruction Algorithm for Automated Microscopic Urinalysis

Microscopic urinalysis is a routine diagnostic test at hospitals. Recent studies have demonstrated the effectiveness of deep learning methods to autom...

Jul 15 2026 2607.13601v1
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-langu...

Jul 15 2026 2607.13931v1
Screening Is Effective for Visual Recognition

Vision Transformer (ViT) has been widely used as a powerful framework for modeling global dependencies among image patches. However, its core componen...

Jul 15 2026 2607.13983v1
Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop cl...

Jul 14 2026 2607.12818v2
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

We present X-lens, a compact feed-forward model for metric depth estimation from a variable number of calibrated fisheye and pinhole views. To support...

Jul 14 2026 2607.12993v2
Self-Supervised Visual Representation Learning: Pretrain-Finetuning or Joint Training?

Self-supervision is a powerful technique for learning visual representations from unlabeled data. Existing techniques primarily adopt a two-stage appr...

Jul 14 2026 2607.13192v1
Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction

Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs o...

Jul 14 2026 2607.12293v1
DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial co...

Jul 14 2026 2607.12319v1
MQAdapter: Multi-Modal Quantum Adapter for Coarse-to-Fine VLM Fine-tuning

Large-scale Vision-Language Models have demonstrated impressive transfer learning capabilities across a wide range of tasks. For few-shot classificati...

Jul 14 2026 2607.12418v1
Let RGB Be the Language of Vision

This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth ma...

Jul 14 2026 2607.12450v1
Towards Vision-Free CIR: Attribute-Augmented Scoring and LLM-Based Reranking for Zero-Shot Composed Image Retrieval

Recent work has shown that "Vision-Free'' approaches (representing images as text) can be effective for standard image retrieval tasks. However, it re...

Jul 14 2026 2607.12621v1
Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop cl...

Jul 14 2026 2607.12818v1
ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models

LiDAR-based collaborative 3D perception in Vehicle-to-Everything (V2X) systems typically relies on fusing bird's-eye-view (BEV) features across agents...

Jul 14 2026 2607.12959v1
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

We present X-lens, a compact feed-forward model for metric depth estimation from a variable number of calibrated fisheye and pinhole views. To support...

Jul 14 2026 2607.12993v1
Enabling 24-hour Agricultural Robotics: Unsupervised Day-to-Night Cross-Modal Image Translation for Nighttime Visual Navigation

While visual navigation has been extensively studied in agricultural robotics, most existing systems assume daytime conditions. In fact, deploying aut...

Jul 13 2026 2607.12065v1
Aqueous Humor Liquid Biopsy Enables Multi-Omics Tumor Profiling and Methylation-Based Machine-Learning Stratification of Retinoblastoma

Primary tumor biopsy in retinoblastoma carries an unacceptable risk of extraocular dissemination. As a result, children treated with eye-sparing appro...

Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools

Recent multimodal large language models (MLLMs) have made remarkable progress on fine-grained perception tasks under the "Thinking with Images" (TwI) ...

Jul 13 2026 2607.11106v1
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation,...

Jul 13 2026 2607.11562v1
Browse Categories