Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4041-4060 of 9,853 articles

SonoCLIP: Mask-Guided Region-Aware Vision-Language Pretraining for Fetal Ultrasound Analysis

Vision-language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging have recently emerged, the domain remains particularly challenging due to severe speckle noise, acquisition variability, and subtle anatomical boundaries, leading to high inter-observer variability. Existing CLIP-based models rely primarily on global image-text...

Jun 28 2026 2606.29586v1

VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models

Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correction, and pixel fidelity. However, such optics are often impractical for size-, cost-, or form-factor-constrained applications, where compact meta-optics offer an attract...

Jun 26 2026 2606.27646v1
A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of $\mathrm{O}(2)$

Vision transformers have become a dominant architecture for visual recognition. However, standard models do not explicitly encode the planar symmetrie...

Jun 26 2026 2606.27864v1
A Multi-Attribute Latent Space for Visual Analysis of Watches

We present a design rationale, embedding model, and interactive visual-analysis system for exploring large wristwatch collections through heterogeneou...

Jun 26 2026 2606.27897v1
Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training

Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where acti...

Jun 26 2026 2606.28104v1
Clinically aligned rationale generation for glaucoma subtype classification via a knowledge-distilled language model

Automated glaucoma subtype classification from clinical notes remains clinically unactionable without subspecialty-aligned explanations supporting cli...

Aloe-Vision: Robust Vision-Language Models for Healthcare

Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinica...

Jun 25 2026 2606.27500v1
Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge

Large Language Models (LLMs) possess broad conceptual knowledge acquired through large-scale text pretraining, yet their potential to supervise models...

Jun 25 2026 2606.27527v1
Can Demographic Information Be Reduced in Retinal Fundus Images While Preserving Glaucoma-Relevant Features?

Purpose: To determine whether disease-aware adversarial perturbations can reduce demographic recoverability encoded in color fundus photographs (CFPs)...

Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks

Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain. We study \textbf{r...

Jun 25 2026 2606.26476v1
TinyCNNDeep: Lightweight Attention-Based CNN for EEG Classification of Eye States and Sleep Deprivation

Sleep deprivation impairs vigilance and cognitive function, yet jointly identifying the sleep condition (normal vs deprived) and the eye state (open v...

Jun 25 2026 2606.26506v1
LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen p...

Jun 25 2026 2606.27192v1
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, rep...

Jun 25 2026 2606.27313v1
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. Howeve...

Jun 25 2026 2606.27373v1
Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models

Visual prompt tuning has emerged as a parameter-efficient fine-tuning approach for adapting large-scale Vision Transformers (ViTs) to downstream tasks...

Jun 24 2026 2606.26379v1
Multilingual Hematology Visual Question Answering Dataset

Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for ...

Jun 24 2026 2606.25246v1
REViT: Roto-reflection Equivariant Convolutional Vision Transformer

In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant ne...

Jun 24 2026 2606.25318v1
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity

Vision Language Models have achieved near-human performance on single-document Visual Question Answering, yet their effectiveness degrades significant...

Jun 24 2026 2606.25343v1
Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet...

Jun 24 2026 2606.25546v1
What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment

Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challengin...

Jun 24 2026 2606.25718v1
Browse Categories