Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4521-4540 of 9,853 articles

V-Bridge: Bridging Video Generative Priors to Versatile Few-shot Image Restoration

Large-scale video generative models are trained on vast and diverse visual data, enabling them to internalize rich structural, semantic, and dynamic priors of the visual world. While these models have demonstrated impressive generative capability, their potential as general-purpose visual learners remains largely untapped. In this work, we introduce V-Bridge, a framework that bridges this latent c...

Mar 13 2026 2603.13089v1

Visual-ERM: Reward Modeling for Visual Equivalence

Vision-to-code tasks require models to reconstruct structured visual inputs, such as charts, tables, and SVGs, into executable or structured representations with high visual fidelity. While recent Large Vision Language Models (LVLMs) achieve strong results via supervised fine-tuning, reinforcement learning remains challenging due to misaligned reward signals. Existing rewards either rely on textua...

Mar 13 2026 2603.13224v1
TumorCLIP: Lightweight Vision-Language Fusion for Explainable MRI-Based Brain Tumor Classification

Accurate classification of brain tumors from MRI is critical for guiding clinical decision-making; however, existing deep learning models are often hi...

Human Knowledge Integrated Multi-modal Learning for Single Source Domain Generalization

Generalizing image classification across domains remains challenging in critical tasks such as fundus image-based diabetic retinopathy (DR) grading an...

Mar 12 2026 2603.12369v1
Deployment-Oriented Session-wise Meta-Calibration for Landmark-Based Webcam Gaze Tracking

Practical webcam gaze tracking is constrained not only by error, but also by calibration burden, robustness to head motion and session drift, runtime ...

Mar 12 2026 2603.12388v1
BackdoorIDS: Zero-shot Backdoor Detection for Pretrained Vision Encoder

Self-supervised and multimodal vision encoders learn strong visual representations that are widely adopted in downstream vision tasks and large vision...

Mar 12 2026 2603.11664v1
Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysis

Prevalent Computational Aberration Correction (CAC) methods are typically tailored to specific optical systems, leading to poor generalization and lab...

Mar 12 2026 2603.12083v1
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams

Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. Howev...

Mar 12 2026 2603.12265v1
Evidential learning driven Breast Tumor Segmentation with Stage-divided Vision-Language Interaction

Breast cancer is one of the most common causes of death among women worldwide, with millions of fatalities annually. Magnetic Resonance Imaging (MRI) ...

Mar 11 2026 2603.11206v1
One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs Hallucination

Current training-free methods tackle MLLM hallucination with separate strategies: either enhancing visual signals or suppressing text inertia. However...

Mar 11 2026 2603.10360v1
Fighting Hallucinations with Counterfactuals: Diffusion-Guided Perturbations for LVLM Hallucination Suppression

While large vision-language models (LVLMs) achieve strong performance on multimodal tasks, they frequently generate hallucinations -- unfaithful outpu...

Mar 11 2026 2603.10470v1
HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement

Generative models are widely employed to enhance the photorealism of synthetic data for training computer vision algorithms. However, they often intro...

Mar 11 2026 2603.10604v1
VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs

Vision-language pretraining has driven significant progress in medical image analysis. However, current methods typically supervise visual encoders us...

Mar 10 2026 2603.09109v2
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity

Vision-language models (VLMs) face significant computational inefficiencies caused by excessive generation of visual tokens. While prior work shows th...

Mar 10 2026 2603.09480v2
A Saccade-inspired Approach to Image Classification using Vision Transformer Attention Maps

Human vision achieves remarkable perceptual performance while operating under strict metabolic constraints. A key ingredient is the selective attentio...

Mar 10 2026 2603.09613v2
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering

Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA ...

Mar 10 2026 2603.09689v2
Can Parents and Patients Understand Myopia Using Large Language Model-Based Chatbots?

Purpose: This study aimed to compare the reliability of myopia-related information from AI chatbots using a set of commonly asked questions by parents...

MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration

Medical vision-language pretraining (VLP) models have recently been investigated for their generalization to diverse downstream tasks. However, curren...

Mar 10 2026 2603.09101v1
VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs

Vision-language pretraining has driven significant progress in medical image analysis. However, current methods typically supervise visual encoders us...

Mar 10 2026 2603.09109v1
Rotation Equivariant Mamba for Vision Tasks

Rotation equivariance constitutes one of the most general and crucial structural priors for visual data, yet it remains notably absent from current Ma...

Mar 10 2026 2603.09138v1
Browse Categories