Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4681-4700 of 9,853 articles

Index Light, Reason Deep: Deferred Visual Ingestion for Visual-Dense Document Question Answering

Existing multimodal document question answering methods universally adopt a supply-side ingestion strategy: running a Vision-Language Model (VLM) on every page during indexing to generate comprehensive descriptions, then answering questions through text retrieval. However, this "pre-ingestion" approach is costly (a 113-page engineering drawing package requires approximately 80,000 VLM tokens), end...

Feb 15 2026 2602.14162v1

Towards Spatial Transcriptomics-driven Pathology Foundation Models

Spatial transcriptomics (ST) provides spatially resolved measurements of gene expression, enabling characterization of the molecular landscape of human tissue beyond histological assessment as well as localized readouts that can be aligned with morphology. Concurrently, the success of multimodal foundation models that integrate vision with complementary modalities suggests that morphomolecular cou...

Feb 15 2026 2602.14177v1
Detection-Guided Artifact Removal for Clinical EEG: A Deep Learning Framework

Objective: We developed and validated a detection-guided artifact removal framework for clinical electroencephalography (EEG). The framework applies a...

Thinking Like a Radiologist: A Dataset for Anatomy-Guided Interleaved Vision Language Reasoning in Chest X-ray Interpretation

Radiological diagnosis is a perceptual process in which careful visual inspection and language reasoning are repeatedly interleaved. Most medical larg...

Feb 13 2026 2602.12843v1
Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues

Recent generative models produce near-photorealistic images, challenging the trustworthiness of photographs. Synthetic image detection (SID) has thus ...

Feb 12 2026 2602.12381v1
A Large Language Model for Disaster Structural Reconnaissance Summarization

Artificial Intelligence (AI)-aided vision-based Structural Health Monitoring (SHM) has emerged as an effective approach for monitoring and assessing s...

Feb 12 2026 2602.11588v1
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models

Recent vision-language-action (VLA) models built upon pretrained vision-language models (VLMs) have achieved significant improvements in robotic manip...

Feb 12 2026 2602.11832v1
Can Local Vision-Language Models improve Activity Recognition over Vision Transformers? -- Case Study on Newborn Resuscitation

Accurate documentation of newborn resuscitation is essential for quality improvement and adherence to clinical guidelines, yet remains underutilized i...

Feb 12 2026 2602.12002v1
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration

Self-play has enabled large language models to autonomously improve through self-generated challenges. However, existing self-play methods for vision-...

Feb 11 2026 2602.11241v1
Chatting with Images for Introspective Visual Thinking

Current large vision-language models (LVLMs) typically rely on text-only reasoning based on a single-pass visual encoding, which often leads to loss o...

Feb 11 2026 2602.11073v2
1%>100%: High-Efficiency Visual Adapter with Complex Linear Projection Optimization

Deploying vision foundation models typically relies on efficient adaptation strategies, whereas conventional full fine-tuning suffers from prohibitive...

Feb 11 2026 2602.10513v1
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This p...

Feb 11 2026 2602.10809v1
Chatting with Images for Introspective Visual Thinking

Current large vision-language models (LVLMs) typically rely on text-only reasoning based on a single-pass visual encoding, which often leads to loss o...

Feb 11 2026 2602.11073v1
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is i...

Feb 10 2026 2602.10179v1
ICODEN: Ordinary Differential Equation Neural Networks for Interval-Censored Data

Predicting time-to-event outcomes when event times are interval censored is challenging because the exact event time is unobserved. Many existing surv...

Feb 10 2026 2602.10303v1
Uncertainty-Aware Ordinal Deep Learning for cross-Dataset Diabetic Retinopathy Grading

Diabetes mellitus is a chronic metabolic disorder characterized by persistent hyperglycemia due to insufficient insulin production or impaired insulin...

Feb 10 2026 2602.10315v1
Understanding and Enhancing Encoder-based Adversarial Transferability against Large Vision-Language Models

Large vision-language models (LVLMs) have achieved impressive success across multimodal tasks, but their reliance on visual inputs exposes them to sig...

Feb 10 2026 2602.09431v1
Delving into Spectral Clustering with Vision-Language Representations

Spectral clustering is known as a powerful technique in unsupervised data analysis. The vast majority of approaches to spectral clustering are driven ...

Feb 10 2026 2602.09586v1
Code2World: A GUI World Model via Renderable Code Generation

Autonomous GUI agents interact with environments by perceiving interfaces and executing actions. As a virtual sandbox, the GUI World model empowers ag...

Feb 10 2026 2602.09856v1
Barycentric alignment for instance-level comparison of neural representations

Comparing representations across neural networks is challenging because representations admit symmetries, such as arbitrary reordering of units or rot...

Feb 9 2026 2602.09225v1
Browse Categories