Latest AI and machine learning research in ophthalmology for healthcare professionals.
Robots struggle to understand object properties like shape, material, and semantics due to limited prior knowledge, hindering manipulation in unstructured environments. In contrast, humans learn these properties through interactive multi-sensor exploration. This work proposes fusing visual and tactile observations into a unified Gaussian Process Distance Field (GPDF) representation for active pe...
Multimodal ophthalmic imaging-based diagnosis integrates color fundus image with optical coherence tomography (OCT) to provide a comprehensive view of ocular pathologies. However, the uneven global distribution of healthcare resources often results in real-world clinical scenarios encountering incomplete multimodal data, which significantly compromises diagnostic accuracy. Existing commonly used...
As a long-term complication of diabetes, diabetic retinopathy (DR) progresses slowly, potentially taking years to threaten vision. An accurate and r...
Synthetic images are an option for augmenting limited medical imaging datasets to improve the performance of various machine learning models. A comm...
We present Saliency Benchmark (SalBench), a novel benchmark designed to assess the capability of Large Vision-Language Models (LVLM) in detecting vi...
Effective cross-modal retrieval is essential for applications like information retrieval and recommendation systems, particularly in specialized dom...
Survival prediction using whole-slide images (WSIs) is crucial in cancer re-search. Despite notable success, existing approaches are limited by thei...
Wallets are access points for the digital economys value creation. Wallets for blockchains store the end-users cryptographic keys for administrating...
Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and const...
Pixels in image sensors have progressively become smaller, driven by the goal of producing higher-resolution imagery. However, ceteris paribus, a sm...
Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore...
We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are lim...
The global demand for radiologists is increasing rapidly due to a growing reliance on medical imaging services, while the supply of radiologists is ...
With the increasing attention to pre-trained vision-language models (VLMs), \eg, CLIP, substantial efforts have been devoted to many downstream task...
Molecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. W...
Despite years of research and the dramatic scaling of artificial intelligence (AI) systems, a striking misalignment between artificial and human vis...
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in interpreting images using natural language. However, without u...
Federated learning (FL) mechanisms typically require each client to transfer their weights to a central server, irrespective of how useful they are....
Medical image recognition serves as a key way to aid in clinical diagnosis, enabling more accurate and timely identification of diseases and abnorma...
Test-Time Adaptation (TTA) has emerged as a promising solution for adapting a source model to unseen medical sites using unlabeled test data, due to...