Latest AI and machine learning research in medicare for healthcare professionals.
Classifying images with an interpretable decision-making process is a long-standing problem in computer vision. In recent years, Prototypical Part Networks has gained traction as an approach for self-explainable neural networks, due to their ability to mimic human visual reasoning by providing explanations based on prototypical object parts. However, the quality of the explanations generated by ...
Long video generation remains a challenging and compelling topic in computer vision. Diffusion based models, among the various approaches to video generation, have achieved state of the art quality with their iterative denoising procedures. However, the intrinsic complexity of the video domain renders the training of such diffusion models exceedingly expensive in terms of both data curation and ...
Visual instructions for long-horizon tasks are crucial as they intuitively clarify complex concepts and enhance retention across extended steps. Dir...
Methods to quantify uncertainty in predictions from arbitrary models are in demand in high-stakes domains like medicine and finance. Conformal predi...
In personalized technology and psychological research, precisely detecting demographic features and personality traits from digital interactions bec...
Long Video Question Answering (LVQA) is challenging due to the need for temporal reasoning and large-scale multimodal data processing. Existing meth...
From image to video understanding, the capabilities of Multi-modal LLMs (MLLMs) are increasingly powerful. However, most existing video understandin...
Smart factories enhance production efficiency and sustainability, but emergencies like human errors, machinery failures and natural disasters pose s...
Large language models (LLMs) are increasingly adopted in medical question-answering (QA) scenarios. However, LLMs can generate hallucinations and no...
Imitation learning frameworks for robotic manipulation have drawn attention in the recent development of language model grounded robotics. However, ...
We address the long-horizon mapless navigation problem: enabling robots to traverse novel environments without relying on high-definition maps or pr...
Diffusion MRI tractography technique enables non-invasive visualization of the white matter pathways in the brain. It plays a crucial role in neuros...
As research in the Scientometric deepens, the impact of data quality on research outcomes has garnered increasing attention. This study, based on We...
Advancing AI in computational pathology requires large, high-quality, and diverse datasets, yet existing public datasets are often limited in organ ...
Generating knowledge-intensive and comprehensive long texts, such as encyclopedia articles, remains significant challenges for Large Language Models...
Online platforms like Pinterest hosting vast content collections traditionally rely on manual curation or user-generated search logs to create keywo...
Precise Event Spotting (PES) aims to identify events and their class from long, untrimmed videos, particularly in sports. The main objective of PES ...
Uncertainty quantification is necessary for developers, physicians, and regulatory agencies to build trust in machine learning predictors and improv...
Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language ...
Quantum federated learning (QFL) merges the privacy advantages of federated systems with the computational potential of quantum neural networks (QNN...