Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3701-3720 of 9,853 articles

Uncertainty of Vision Medical Foundation Models

Accurate uncertainty estimation is essential for machine learning systems de- ployed in high-stakes domains such as medicine. Traditional approaches primarily rely on probability outputs from trained models (point predictions), which provide no formal guarantees on prediction coverage and often require additional calibra- tion techniques to improve reliability. In contrast, conformal prediction (r...

Aug 31 2026 2608.30390v1

VisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMs

Object hallucination remains a persistent reliability issue in large vision-language models, where generated object mentions may sound plausible but lack visual grounding. Recent training-free detectors use internal signals such as token likelihood, attention, visual confidence, or image-text similarity to identify hallucinated objects. These signals are useful, but they are often source-confounde...

Aug 31 2026 2608.30480v1
Learning Dynamics of Logits Debiasing for Long-Tailed Semi-Supervised Learning

Long-tailed distributions are prevalent in real-world semi-supervised learning (SSL), where pseudo-labels tend to favor majority classes, leading to d...

Aug 31 2026 2608.30699v1
VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs

Multimodal large language models (MLLMs) struggle with fine-grained Visual Search, the task of locating small or rare objects in high-resolution image...

Aug 31 2026 2608.30705v1
Do VLMs Share Safety Neurons Across Modalities?

Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content i...

Aug 31 2026 2608.30750v1
A Composition-Aware Pretraining Framework for Geospatial Foundation Models

Geospatial foundation models have emerged as state-of-the-art methods for downstream Earth observation tasks. However, existing pretraining methodolog...

Aug 31 2026 2608.30817v1
LOCI: A Locator-Critic with Refinement Loop

Vision-Language Models (VLMs) still struggle on tasks requiring complex visual understanding. We argue that the core issue is not high-level reasoning...

Aug 31 2026 2608.30959v1
Robust retinal biometrics for patient identity verification and retrieval across age and imaging devices

Patient identity errors can compromise longitudinal medical records, research databases, and downstream clinical decisions. We present a retinal biome...

Aug 31 2026 2608.31094v1
Seeing Through Extreme Visual Sparsity: Surface Understanding from a Single Random Visual Patch

Surface material recognition from incomplete visual observations remains a challenging problem in robotic perception and environmental understanding. ...

Aug 30 2026 2608.29475v1
Conducting Stylistic Analysis of Paintings through an Art-History Agent

Attributing an artwork to an artist has traditionally relied on detailed visual observations and descriptions, known as stylistic analysis in art hist...

Aug 30 2026 2608.29644v1
Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models

Large Vision-Language Models (LVLMs) achieve strong performance across many multimodal tasks; however, they often exploit spurious object-background c...

Aug 30 2026 2608.29996v1
GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception

Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for sce...

Aug 28 2026 2608.27971v1
Explainable Diabetic Retinopathy Classification Using Vision Foundation Models

Diabetic retinopathy (DR) is a major cause of preventable blindness, creating a need for accurate and trustworthy automated screening. This study inve...

Aug 28 2026 2608.28207v1
AIM: Anchor Identity Features, Then Match for Multimodal Large Language Model Unlearning

Multimodal large language models (MLLMs) can memorize identity-specific facts about people in their fine-tuning data, creating privacy risks when a pe...

Aug 28 2026 2608.28312v1
ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT

Contrastive vision-language learning uses paired chest CT volumes and radiology reports to learn abnormality classifiers without manually annotated la...

Aug 28 2026 2608.28455v1
Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning

Many multimodal tasks depend on how visual elements are ordered and composed, not only on recognizing them in isolation. Internet memes are a compact ...

Aug 27 2026 2608.26866v2
DIRECT: Decomposing Audience Preference and Creative Effect in Visual Content Analytics

Which visual choices make a post perform better? A growing literature answers this question with pooled coefficients estimated across many creators, w...

Aug 27 2026 2608.26584v1
Multi-Image Visual Token Pruning in Large Visual Language Models

With the growing demand for processing multiple image sequences in real-world applications, various visual token pruning methods have emerged to mitig...

Aug 27 2026 2608.26806v1
Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning

Many multimodal tasks depend on how visual elements are ordered and composed, not only on recognizing them in isolation. Internet memes are a compact ...

Aug 27 2026 2608.26866v1
Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations

Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect visual data while securely processing i...

Aug 27 2026 2608.27066v1
Browse Categories