Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4721-4740 of 9,853 articles

VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers

Replicating In-Context Learning (ICL) in computer vision remains challenging due to task heterogeneity. We propose \textbf{VIRAL}, a framework that elicits visual reasoning from a pre-trained image editing model by formulating ICL as conditional generation via visual analogy ($x_s : x_t :: x_q : y_q$). We adapt a frozen Diffusion Transformer (DiT) using role-aware multi-image conditioning and intr...

Feb 3 2026 2602.03210v1

Full end-to-end diagnostic workflow automation of 3D OCT via foundation model-driven AI for retinal diseases

Optical coherence tomography (OCT) has revolutionized retinal disease diagnosis with its high-resolution and three-dimensional imaging nature, yet its full diagnostic automation in clinical practices remains constrained by multi-stage workflows and conventional single-slice single-task AI models. We present Full-process OCT-based Clinical Utility System (FOCUS), a foundation model-driven framework...

Feb 3 2026 2602.03302v1
Origin Lens: A Privacy-First Mobile Framework for Cryptographic Image Provenance and AI Detection

The proliferation of generative AI poses challenges for information integrity assurance, requiring systems that connect model governance with end-user...

Feb 3 2026 2602.03423v1
Contextualized Visual Personalization in Vision-Language Models

Despite recent progress in vision-language models (VLMs), existing approaches often fail to generate personalized responses based on the user's specif...

Feb 3 2026 2602.03454v1
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs

Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a vid...

Feb 3 2026 2602.03615v1
Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis

Retinal diseases spanning a broad spectrum can be effectively identified and diagnosed using complementary signals from multimodal data. However, mult...

Feb 3 2026 2602.03622v1
FOVI: A biologically-inspired foveated interface for deep vision models

Human vision is foveated, with variable resolution peaking at the center of a large field of view; this reflects an efficient trade-off for active sen...

Feb 3 2026 2602.03766v1
End-to-end reconstruction of OCT optical properties and speckle-reduced structural intensity via physics-based learning

Inverse scattering in optical coherence tomography (OCT) seeks to recover both structural images and intrinsic tissue optical properties, including re...

Feb 2 2026 2602.02721v1
Causality--Δ: Jacobian-Based Dependency Analysis in Flow Matching Models

Flow matching learns a velocity field that transports a base distribution to data. We study how small latent perturbations propagate through these flo...

Feb 2 2026 2602.02793v1
Neural Networks as Entropic Systems: Applications in Digital Pathology

Deep learning systems in digital pathology are widely regarded as opaque, limiting clinical trust and interpretability. We present a framework for emp...

Toward a Machine Bertin: Why Visualization Needs Design Principles for Machine Cognition

Visualization's design knowledge-effectiveness rankings, encoding guidelines, color models, preattentive processing rules -- derives from six decades ...

Feb 2 2026 2602.01527v1
Preserving Localized Patch Semantics in VLMs

Logit Lens has been proposed for visualizing tokens that contribute most to LLM answers. Recently, Logit Lens was also shown to be applicable in autor...

Feb 2 2026 2602.01530v1
SGHA-Attack: Semantic-Guided Hierarchical Alignment for Transferable Targeted Attacks on Vision-Language Models

Large vision-language models (VLMs) are vulnerable to transfer-based adversarial perturbations, enabling attackers to optimize on surrogate models and...

Feb 2 2026 2602.01574v1
Federated Vision Transformer with Adaptive Focal Loss for Medical Image Classification

While deep learning models like Vision Transformer (ViT) have achieved significant advances, they typically require large datasets. With data privacy ...

Feb 2 2026 2602.01633v1
Efficient Cross-Country Data Acquisition Strategy for ADAS via Street-View Imagery

Deploying ADAS and ADS across countries remains challenging due to differences in legislation, traffic infrastructure, and visual conventions, which i...

Feb 2 2026 2602.01836v1
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning

Large multimodal reasoning models solve challenging visual problems via explicit long-chain inference: they gather visual clues from images and decode...

Feb 2 2026 2602.02004v1
Rethinking Genomic Modeling Through Optical Character Recognition

Recent genomic foundation models largely adopt large language model architectures that treat DNA as a one-dimensional token sequence. However, exhaust...

Feb 2 2026 2602.02014v1
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models

Modern Vision-Language Models (VLMs) exhibit a critical flaw in compositional reasoning, often confusing "a red cube and a blue sphere" with "a blue c...

Feb 2 2026 2602.02043v1
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have advanced VQA and now support Vision-DeepResearch systems that use search engines for complex visual-text...

Feb 2 2026 2602.02185v1
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models

The integration of visual information into Large Language Models (LLMs) has enabled Multimodal LLMs (MLLMs), but the quadratic memory and computationa...

Feb 2 2026 2602.02197v1
Browse Categories