Latest AI and machine learning research in ophthalmology for healthcare professionals.
The rising prevalence of autism spectrum disorder (ASD) strains clinical infrastructure. Gold-standard tools like ADOS-2 face high costs, specialized training requirements, and extensive waitlists, delaying diagnosis and intervention. While eye-tracking offers a promising digital biomarker, existing tools lack scalable community deployment due to hardware costs and operational constraints. Here, w...
We investigate the token mixer in vision backbones by revisiting clustering, one of the most classic approaches in machine learning. An effective token mixer is a fundamental component of modern vision backbones like vision Transformers, facilitating information exchange between image patches. Mainstream token mixers, which rely on convolution, attention, MLP, or their hybrids, primarily focus on ...
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution ...
Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory overhead. Exis...
Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal ...
Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete cras...
Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understandin...
Chimeric antigen receptor (CAR) cell therapy has achieved transformative clinical success through targeting of CD19 in refractory B cell malignancies,...
Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods ado...
We propose a novel object detection method that enables us to protect sensitive visual information of test images. Previous studies considering visual...
Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, ...
Objective: Cochlear implants (CIs) are bionic prostheses that restores hearing via electrical stimulation of the auditory nerve. Hybrid CIs, which use...
Leading deep neural network encoding models predict visual cortical responses with nearly indistinguishable accuracy, raising the strong inference tha...
Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI s...
High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured b...
Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual e...
Electroretinography (ERG) measures the functional response of distinct retinal cells to light, but was largely displaced by structural imaging in the ...
Precise segmentation of the optic disc and cup is critical for the early detection and diagnosis of glaucoma. However, achieving consistently high per...
Diffusion-based generators have made synthetic images ubiquitous, but detectors often fail under simultaneous shifts in generator, prompt/style, and s...
Aligned vision-language models (VLMs) are designed to balance grounded visual reasoning with safe generation behavior. However, we observe a striking ...