Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4701-4720 of 9,853 articles

CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT

Recent advances in vision-language models (VLMs) have markedly improved image-text alignment, yet they still fall short of human-like visual reasoning. A key limitation is that many VLMs rely on surface correlations rather than building logically coherent structured representations, which often leads to missed higher-level semantic structure and non-causal relational understanding, hindering compo...

Feb 9 2026 2602.08339v1

What, Whether and How? Unveiling Process Reward Models for Thinking with Images Reasoning

The rapid advancement of Large Vision Language Models (LVLMs) has demonstrated excellent abilities in various visual tasks. Building upon these developments, the thinking with images paradigm has emerged, enabling models to dynamically edit and re-encode visual information at each reasoning step, mirroring human visual processing. However, this paradigm introduces significant challenges as diverse...

Feb 9 2026 2602.08346v1
OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Hypothesis. Artificial general intelligence is, at its core, a compression problem. Effective compression demands resonance: deep learning scales best...

Feb 9 2026 2602.08683v1
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks

Vision-Language Models (VLMs) are now a core part of modern AI. Recent work proposed several visual jailbreak attacks using single/ holistic images. H...

Feb 8 2026 2602.08136v1
Accelerating Vision Transformers on Brain Processing Unit

With the advancement of deep learning technologies, specialized neural processing hardware such as Brain Processing Units (BPUs) have emerged as dedic...

Feb 6 2026 2602.06300v1
A neuromorphic model of the insect visual system for natural image processing

Insect vision supports complex behaviors including associative learning, navigation, and object detection, and has long motivated computational models...

Feb 6 2026 2602.06405v1
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs

Despite recent successes, test-time scaling - i.e., dynamically expanding the token budget during inference as needed - remains brittle for vision-lan...

Feb 6 2026 2602.06566v1
MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

Multimodal large language models (MLLMs) have rapidly advanced, yet their adoption in medicine remains limited by gaps in domain coverage, modality al...

Feb 6 2026 2602.06965v1
VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs

Recent progress in Vision Language Models (VLMs) has raised the question of whether they can reliably perform nonverbal reasoning. To this end, we int...

Feb 5 2026 2602.05382v1
Multi-AD: Cross-Domain Unsupervised Anomaly Detection for Medical and Industrial Applications

Traditional deep learning models often lack annotated data, especially in cross-domain applications such as anomaly detection, which is critical for e...

Feb 5 2026 2602.05426v1
Visual Implicit Geometry Transformer for Autonomous Driving

We introduce the Visual Implicit Geometry Transformer (ViGT), an autonomous driving geometric model that estimates continuous 3D occupancy fields from...

Feb 5 2026 2602.05573v1
Unseen Insights: An AI-Powered Exploration of Secure Patient Messages in Ophthalmology

Objective To characterize the clinical and administrative concerns communicated through secure ophthalmology messaging and to assess differences in me...

Live-cell imaging enables reporter-free monitoring of the circadian rhythm in individual Synechocystis cells

In vivo monitoring of circadian rhythms depends on reliable and non-invasive detection methods. This is often achieved by expressing reporter genes he...

Partial Ring Scan: Revisiting Scan Order in Vision State Space Models

State Space Models (SSMs) have emerged as efficient alternatives to attention for vision tasks, offering lineartime sequence processing with competiti...

Feb 4 2026 2602.04170v1
Beyond Static Cropping: Layer-Adaptive Visual Localization and Decoding Enhancement

Large Vision-Language Models (LVLMs) have advanced rapidly by aligning visual patches with the text embedding space, but a fixed visual-token budget f...

Feb 4 2026 2602.04304v1
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning

Large foundation models have shown strong open-world generalization to complex problems in vision and language, but similar levels of generalization h...

Feb 4 2026 2602.04315v1
Understanding Degradation with Vision Language Model

Understanding visual degradations is a critical yet challenging problem in computer vision. While recent Vision-Language Models (VLMs) excel at qualit...

Feb 4 2026 2602.04565v1
When LLaVA Meets Objects: Token Composition for Vision-Language-Models

Current autoregressive Vision Language Models (VLMs) usually rely on a large number of visual tokens to represent images, resulting in a need for more...

Feb 4 2026 2602.04864v1
Deep-learning-based fMRI decoding of real-world size for hand-held objects

Real-world object size is a fundamental dimension of visual cognition, supporting effective interaction with the environment and object manipulation. ...

Aligning Forest and Trees in Images and Long Captions for Visually Grounded Understanding

Large vision-language models such as CLIP struggle with long captions because they align images and texts as undifferentiated wholes. Fine-grained vis...

Feb 3 2026 2602.02977v1
Browse Categories