Latest AI and machine learning research in ophthalmology for healthcare professionals.
Abstract Metagenomic sequencing can detect a broad range of pathogens, but interpreting which detections are clinically relevant requires expert adjudication that is difficult to scale and standardize. Here we present diagnostic classifiers that formalize expert adjudication by combining structured decision trees with large language model reasoning to assign diagnoses and select pathogen candidate...
As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades, yet routine assessment often misses its earliest signs. This article critically synthesizes recent technological advances for detecting and managing cognitive impairment in older adults, spanning neurophysiological signals (chiefly electroencephalography, EEG...
Large vision-language models (LVLMs) often hallucinate objects that are absent from an image. Despite recent progress, existing mitigation methods sti...
Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged a...
Monocular depth estimation (MDE) faces challenges with non-Lambertian surfaces and adverse weather conditions due to the visual ambiguities inherent i...
Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens a...
Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capa...
Medical image segmentation heavily depends on labor-intensive and time-consuming pixel-level annotations. Eye tracking offers a cost-effective solutio...
Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images ge...
Anterior eye segment (AES) segmentation is a key component of both ocular biometrics and emerging clinical image analysis applications. However, heter...
CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume repr...
Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associatin...
Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. H...
Background: Automated behavioral tracking is increasingly used in biological and biomedical research; however, robustness across heterogeneous imaging...
Glaucoma is a group of eye diseases that damage the optic nerve, often caused by elevated intraocular pressure. It is a leading cause of irreversible ...
Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction assumes th...
Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data d...
Face de-identification (De-ID) aims to remove or conceal personally identifiable facial features in images or videos to prevent identity recognition w...
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundam...
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streamin...