Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4741-4760 of 9,853 articles

Unified Personalized Reward Model for Vision Generation

Recent advancements in multimodal reward models (RMs) have significantly propelled the development of visual generation. Existing frameworks typically adopt Bradley-Terry-style preference modeling or leverage generative VLMs as judges, and subsequently optimize visual generation models via reinforcement learning. However, current RMs suffer from inherent limitations: they often follow a one-size-f...

Feb 2 2026 2602.02380v1

Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models

Vision-language models (VLMs) extend large language models (LLMs) with vision encoders, enabling text generation conditioned on both images and text. However, this multimodal integration expands the attack surface by exposing the model to image-based jailbreaks crafted to induce harmful responses. Existing gradient-based jailbreak methods transfer poorly, as adversarial patterns overfit to a singl...

Feb 1 2026 2602.01025v1
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance

Large Vision-Language Models (LVLMs) can reason effectively from image-text inputs and perform well in various multimodal tasks. Despite this success,...

Feb 1 2026 2602.01047v1
Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs

Vision-Language-Action (VLA) models have emerged as a dominant paradigm for generalist robotic manipulation, unifying perception and control within a ...

Feb 1 2026 2602.01158v1
Bridging Lexical Ambiguity and Vision: A Mini Review on Visual Word Sense Disambiguation

This paper offers a mini review of Visual Word Sense Disambiguation (VWSD), which is a multimodal extension of traditional Word Sense Disambiguation (...

Feb 1 2026 2602.01193v1
Med3D-R1: Incentivizing Clinical Reasoning in 3D Medical Vision-Language Models for Abnormality Diagnosis

Developing 3D vision-language models with robust clinical reasoning remains a challenge due to the inherent complexity of volumetric medical imaging, ...

Feb 1 2026 2602.01200v1
Where to Attend: A Principled Vision-Centric Position Encoding with Parabolas

We propose Parabolic Position Encoding (PaPE), a parabola-based position encoding for vision modalities in attention-based architectures. Given a set ...

Feb 1 2026 2602.01418v1
Theoretical Analysis of Measure Consistency Regularization for Partially Observed Data

The problem of corrupted data, missing features, or missing modalities continues to plague the modern machine learning landscape. To address this issu...

Feb 1 2026 2602.01437v1
EndoCaver: Handling Fog, Blur and Glare in Endoscopic Images via Joint Deblurring-Segmentation

Endoscopic image analysis is vital for colorectal cancer screening, yet real-world conditions often suffer from lens fogging, motion blur, and specula...

Jan 30 2026 2601.22537v1
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration

Multimodal large language models (MLLMs) suffer from high computational costs due to excessive visual tokens, particularly in high-resolution and vide...

Jan 30 2026 2601.22674v1
Jailbreaks on Vision Language Model via Multimodal Reasoning

Vision-language models (VLMs) have become central to tasks such as visual question answering, image captioning, and text-to-image generation. However,...

Jan 29 2026 2601.22398v1
Machine Learning-Based Prediction of Postoperative Refraction in Cataract Surgery: A Stacking Ensemble Approach

Background: Achieving precise postoperative refractive outcomes remains a significant challenge in cataract surgery. While advanced intraocular lens (...

Spatial mapping of RNA turnover kinetics and regulatory landscapes of mRNA stability in the mammalian brain

Spatial and activity-dependent gene regulation in the mammalian brain requires coordinated control of RNA synthesis and degradation, yet spatially res...

Machine Learning Ensemble Reveals Distinct Molecular Pathways of Retinal Damage in Spaceflown Mice

Spaceflight-associated neuro-ocular syndrome (SANS) threatens astronaut health during long-duration missions, yet its molecular pathology remains uncl...

Do eyes say it all? Assessing the utility of physiological signals for predicting systemic cognitive states

Robust estimation of systemic human cognitive states is critical for many applications, from simply detecting inefficiencies in human task performance...

Do Pathology Foundation Models Encode Disease Progression? A Pseudotime Analysis of Visual Representations

Vision foundation models trained on discretely sampled images achieve strong performance on classification benchmarks, yet whether their representatio...

Jan 29 2026 2601.21334v1
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models

The development of large vision language models drives the demand for managing, and applying massive amounts of multimodal data, making OCR technology...

Jan 29 2026 2601.21639v1
Improving Classifier-Free Guidance of Flow Matching via Manifold Projection

Classifier-free guidance (CFG) is a widely used technique for controllable generation in diffusion and flow-based models. Despite its empirical succes...

Jan 29 2026 2601.21892v1
Just Noticeable Difference Modeling for Deep Visual Features

Deep visual features are increasingly used as the interface in vision systems, motivating the need to describe feature characteristics and control fea...

Jan 29 2026 2601.21933v1
Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models

Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity o...

Jan 29 2026 2601.22060v1
Browse Categories