Latest AI and machine learning research in ophthalmology for healthcare professionals.
A method is presented for accelerating inference in transformer language models by exploiting the low effective rank of the token activation manifold at each layer. The method decomposes each activation vector into a subspace component and a residual, computes the linear-layer output on the subspace component via a cached low-rank weight image at reduced memory bandwidth, and applies a per-token g...
Purpose: Assessing visual function in patients with ultra-low vision (ULV), particularly those with retinitis pigmentosa (RP), remains a significant challenge in therapeutic development. Full-field stimulus test (FST) provides a quantitative measure of retinal light sensitivity and may serve as a valuable clinical endpoint. We investigated FST in ULV RP by examining its associations with functiona...
The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak at...
Effectively stratifying patient risk in chronic diseases like glaucoma is a major clinical challenge. Clinicians need tools to identify patients at hi...
Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computatio...
In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining fra...
As Vision-Language Models (VLMs) become increasingly integrated into decision-making systems, it is essential to understand how visual inputs influenc...
Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current v...
Brain encoding models not only serve to decipher how visual stimuli are transformed into neural responses, but also represent a critical step toward v...
Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current v...
Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language l...
Large Vision-Language Models (VLMs) have achieved remarkable multimodal performance yet remain prone to factual hallucinations, particularly in long-t...
One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack...
Diffusion models have achieved remarkable success in synthesizing complex static and temporal visuals, a breakthrough largely driven by Classifier-Fre...
Aerial-Ground Re-Identification (AG-ReID) is constrained by the viewpoint-domain gap, as drastic viewpoint disparities occlude or distort discriminati...
Despite the rapid progress in data-driven 3D vision, aerial geometric 3D vision remains a formidable challenge due to the severe scarcity of large-sca...
Retinal fluids, detectable through optical coherence tomography (OCT), are key biomarkers for retinal diseases such as diabetic macular edema and age-...
Background: Diabetic retinopathy (DR) is the leading cause of preventable blindness among working-age adults worldwide, yet screening coverage remains...
Despite strong performance of deep learning models in retinal disease detection, most systems produce static predictions without clinical reasoning or...
Vision-Language Models (VLMs) exhibit strong performance in instruction following and open-ended vision-language reasoning, yet they frequently genera...