Latest AI and machine learning research in ophthalmology for healthcare professionals.
Vision-language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging have recently emerged, the domain remains particularly challenging due to severe speckle noise, acquisition variability, and subtle anatomical boundaries, leading to high inter-observer variability. Existing CLIP-based models rely primarily on global image-text...
Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correction, and pixel fidelity. However, such optics are often impractical for size-, cost-, or form-factor-constrained applications, where compact meta-optics offer an attract...
Vision transformers have become a dominant architecture for visual recognition. However, standard models do not explicitly encode the planar symmetrie...
We present a design rationale, embedding model, and interactive visual-analysis system for exploring large wristwatch collections through heterogeneou...
Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where acti...
Automated glaucoma subtype classification from clinical notes remains clinically unactionable without subspecialty-aligned explanations supporting cli...
Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinica...
Large Language Models (LLMs) possess broad conceptual knowledge acquired through large-scale text pretraining, yet their potential to supervise models...
Purpose: To determine whether disease-aware adversarial perturbations can reduce demographic recoverability encoded in color fundus photographs (CFPs)...
Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain. We study \textbf{r...
Sleep deprivation impairs vigilance and cognitive function, yet jointly identifying the sleep condition (normal vs deprived) and the eye state (open v...
The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen p...
A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, rep...
Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. Howeve...
Visual prompt tuning has emerged as a parameter-efficient fine-tuning approach for adapting large-scale Vision Transformers (ViTs) to downstream tasks...
Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for ...
In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant ne...
Vision Language Models have achieved near-human performance on single-document Visual Question Answering, yet their effectiveness degrades significant...
Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet...
Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challengin...