Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3981-4000 of 9,853 articles

InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models

Infrared vision-language models are increasingly used for perception under low-light and adverse visual conditions, yet their robustness to localized structured perturbations remains underexplored. Existing infrared adversarial studies mainly focus on object detectors, leaving the security of infrared vision-language models less systematically examined. We present InfraQR, a QR-inspired structured...

Jul 8 2026 2607.07288v1

Guidance Breaks the Fitted Operator: A Terminal-Fitted Repair for Classifier-Free Guidance

Classifier-free guidance (CFG) is the standard way to strengthen class-conditioning in diffusion and flow-matching samplers, yet at large guidance it oversaturates and destabilizes, symptoms practitioners suppress with more steps or limited-interval schedules. We analyze CFG through an asymptotic-preserving, numerical-analysis lens. Building on a recent result that the deterministic DDIM step is t...

Jul 8 2026 2607.07665v1
Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review

Vision Language Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundation model, a...

Jul 7 2026 2607.06706v1
Generative and discriminative recurrence employ opposing strategies for robust vision

Recurrence is thought to enhance the robustness of biological vision, but how it achieves this feat is largely unknown. Perceptual robustness can be i...

Complementary Roles of Image Classification and Vessel Segmentation in AI-Based Screening for Retinopathy of Prematurity Plus Disease in a Kenyan Preterm Cohort

Background. Retinopathy of prematurity (ROP) is a preventable cause of childhood blindness, with rising burden in low- and middle-income countries whe...

Jul 7 2026 2607.05825v1
Realistic Compound-Lens Defocus Blur Synthesis

Defocus blur degrades fine image structures and limits visual perception, which can adversely affect downstream vision tasks. Although recent deep lea...

Jul 7 2026 2607.05837v1
AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring

Vision-Language Models (VLMs) are promising for construction-site monitoring, and recent construction-tailored VLMs have primarily adapted pretrained ...

Jul 7 2026 2607.05859v1
UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex pro...

Jul 7 2026 2607.06306v1
What Images Cannot Say: Language-Guided Olfactory Representation Learning

Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with electronic-nose ...

Jul 7 2026 2607.06402v1
A VLM-Enhanced Framework for Comprehensive Traffic Sign Condition Assessment Integrating Daytime Visual Performance and Nighttime Retroreflectivity Evaluation

Traffic signs are crucial components of road safety, serving as visual tools under all lighting conditions. The Manual on Uniform Traffic Control Devi...

Jul 7 2026 2607.06478v1
MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation

Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB image...

Jul 7 2026 2607.06552v1
Vision as Unified Multimodal Generation

We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation ...

Jul 7 2026 2607.06560v1
SSA-3DGS: Unsupervised Removal of Screen-Space Artifacts for 3D Gaussian Splatting

Novel View Synthesis (NVS) methods, such as 3D Gaussian Splatting (3DGS), rely heavily on the assumption of clean, multi-view consistent, posed input ...

Jul 6 2026 2607.05598v1
Cross-Contextual Vision-Language Adaptation with LoRA for Personalized Severe Adverse Event Detection in Clinical Wound Monitoring

Wound monitoring is a critical yet underserved clinical challenge, where timely identification of severe adverse events (SAEs) such as infection, tiss...

Jul 6 2026 2607.05625v1
VEIL: How Visual Encoding Hijacking Induces Bias In Vision Models

Rendering time series as chart images for CNN-based classification has become increasingly common in time-series classification (TSC). However, it rem...

Jul 6 2026 2607.05641v1
RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather

Robust 3D object detection in adverse weather conditions is challenging due to sensor limitations. Although combining complementary modalities such as...

Jul 6 2026 2607.04587v1
TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models

Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large...

Jul 6 2026 2607.04593v1
Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens...

Jul 6 2026 2607.04605v1
PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving

Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalizati...

Jul 6 2026 2607.04637v1
Integrating Neural Encoders in Bayesian Generalized Linear Mixed Models for Multimodal Data

Scalable Bayesian inference for generalized linear mixed models (GLMMs) provides uncertainty-aware analysis of correlated longitudinal data, but exist...

Jul 6 2026 2607.04647v1
Browse Categories