Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4221-4240 of 9,853 articles

Urban Risk-Aware Navigation via VQA-Based Event Maps for People with Low Vision

Visual impairment affects hundreds of millions of people worldwide, severely limiting their ability to navigate urban environments safely and independently. While wearable assistive devices offer a promising platform for real-time hazard detection, existing approaches rely on task-specific vision pipelines that lack flexibility and generalizability. In this work, we propose an event map framework ...

May 12 2026 2605.11782v1

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement

Large Vision-Language Models (LVLMs) have achieved remarkable performance on diverse vision-language tasks. However, LVLMs still suffer from hallucinations, generating text that contradicts the visual input. Existing research has primarily focused on mitigating object hallucinations, but often overlooks more complex relation hallucinations, particularly action relations involving interactions betw...

May 12 2026 2605.11808v1
REFNet++: Multi-Task Efficient Fusion of Camera and Radar Sensor Data in Bird's-Eye Polar View

A realistic view of the vehicle's surroundings is generally offered by camera sensors, which is crucial for environmental perception. Affordable radar...

May 12 2026 2605.11824v1
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs

Multimodal large language models are increasingly expected to perform thinking with images, yet existing visual latent reasoning methods still rely on...

May 12 2026 2605.11856v1
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models

Multimodal video summarization requires visual features that align semantically with language generation. Traditional approaches rely on CNN features ...

May 12 2026 2605.11959v1
Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters

Vision Large Language Models (VLLMs) have achieved remarkable success in modern text-rich visual understanding. However, their perceptual robustness i...

May 12 2026 2605.11960v1
DistractMIA: Black-Box Membership Inference on Vision-Language Models via Semantic Distraction

Vision-language models (VLMs) are trained on large-scale image-text corpora that may contain private, copyrighted, or otherwise sensitive data, motiva...

May 12 2026 2605.12574v1
Spectral Vision Transformer for Efficient Tokenization with Limited Data

We propose a novel spectral vision transformer architecture for efficient tokenization in limited data, with an emphasis on medical imaging. We outlin...

May 12 2026 2605.12026v1
MULTI: Disentangling Camera Lens, Sensor, View, and Domain for Novel Image Generation

Recent text-to-image models produce high-quality images, yet text ambiguity hinders precise control when specific styles or objects are required. Ther...

May 12 2026 2605.12134v1
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model

In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likew...

May 12 2026 2605.12163v2
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm

Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual gener...

May 12 2026 2605.12271v1
Large-Small Model Collaboration for Farmland Semantic Change Detection

Farmland Semantic Change Detection (SCD) is essential for cultivated land protection, yet existing benchmarks and models remain insufficient for fine-...

May 12 2026 2605.12282v1
Elastic Attention Cores for Scalable Vision Transformers

Vision Transformers (ViTs) achieve strong data-driven scaling by leveraging all-to-all self-attention. However, this flexibility incurs a computationa...

May 12 2026 2605.12491v1
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as dis...

May 12 2026 2605.12500v1
Enhanced processing of cartoons in infant visual cortex

Developing sensory systems may have heightened sensitivity to exaggerated features that emphasize diagnostic information, as shown by the benefits of ...

Measurement-Adapted Eigentask Representations for Photon-Limited Optical Readout

Optical readout in low-light imaging is fundamentally limited by measurement noise, including photon shot noise, detector noise, and quantization erro...

May 11 2026 2605.10008v1
MFVLR: Multi-domain Fine-grained Vision-Language Reconstruction for Generalizable Diffusion Face Forgery Detection and Localization

The swift advancement in photo-realistic face generation technology has sparked considerable concerns across society and academia, emphasizing the req...

May 11 2026 2605.10071v1
AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation

Visual anomaly detection (VAD) is crucial in many real-world fields, such as industrial inspection, medical imaging, infrastructure monitoring, and re...

May 11 2026 2605.10397v1
SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models

Low-level visual perception underpins reliable remote sensing (RS) image analysis, yet current image quality assessment (IQA) methods output uninterpr...

May 11 2026 2605.10576v1
Polygon-mamba: Retinal vessel segmentation using polygon scanning mamba and space-frequency collaborative attention

Retinal vessel segmentation is crucial for diagnosis and assessment of ocular diseases. Notably, segmentation of small retinal vessels has been consis...

May 11 2026 2605.10581v1
Browse Categories