Latest AI and machine learning research in ophthalmology for healthcare professionals.
Existing research on federated learning (FL) usually assumes that training labels are of high quality for each client, which is impractical in many real-world scenarios (e.g., noisy labels by crowd-sourced annotations), leading to dramatic performance degradation. In this work, we investigate noisy FL through the lens of early-time training phenomenon (ETP). Specifically, a key finding of this pap...
Neural models produce promising results when solving Vehicle Routing Problems (VRPs), but may often fall short in generalization. Recent attempts to enhance model generalization often incur unnecessarily large training cost or cannot be directly applied to other models solving different VRP variants. To address these issues, we take a novel perspective on model architecture in this study. Specific...
Traditional convolutional neural networks (CNNs) struggle with scale and rotation transformations, resulting in reduced performance on transformed ima...
BACKGROUND: Large language models (LLMs) have shown promise in various medical applications, but their potential as decision support tools in emergenc...
Geographic atrophy (GA) secondary to age-related macular degeneration is a chronic degenerative disease involving the retinal pigment epithelium, phot...
Vision-Language Models (vLLMs) have emerged as powerful architectures for joint reasoning over visual and textual inputs, enabling breakthroughs in ...
Enabling robots to understand human gaze target is a crucial step to allow capabilities in downstream tasks, for example, attention estimation and m...
High-resolution imaging is crucial for enhancing visual clarity and enabling precise computer-assisted guidance in minimally invasive surgery (MIS)....
Recent progress in multimodal reasoning has been significantly advanced by textual Chain-of-Thought (CoT), a paradigm where models conduct reasoning...
Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visu...
We propose VisTex-OVLM, a novel image prompted object detection method that introduces visual textualization -- a process that projects a few visual...
In real-world vision-language applications, practitioners increasingly rely on large, pretrained foundation models rather than custom-built solution...
Driven by large-scale contrastive vision-language pre-trained models such as CLIP, recent advancements in the image-text matching task have achieved...
The success of machine learning models in industrial applications is heavily dependent on the quality of the datasets used to train the models. Howe...
Vision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain. We in...
In the rapidly evolving field of optical engineering, precise alignment of multi-lens imaging systems is critical yet challenging, as even minor mis...
The global prevalence of dementia is projected to double by 2050, highlighting the urgent need for scalable diagnostic tools. This study utilizes di...
Large Vision-Language Models (VLMs) have revolutionized computer vision, enabling tasks such as image classification, captioning, and visual questio...
Medical visual question answering (MedVQA) plays a vital role in clinical decision-making by providing contextually rich answers to image-based quer...
This work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule...