Latest AI and machine learning research in care of terminally ill / palliative care for healthcare professionals.
Text-to-Image (T2I) synthesis is a challenging task that requires modeling complex interactions between two modalities ( i.e., text and image). A common framework adopted in recent state-of-the-art approaches to achieving such multimodal interactions is to bootstrap the learning process with pre-trained image-aligned text embeddings trained using contrastive loss. Furthermore, these embeddings a...
End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propose OmniHuman, a Diffusion Transformer-based framework that scales up data by mixing motion-related...
Visual speech recognition remains an open research problem where different challenges must be considered by dispensing with the auditory sense, such...
Though end-to-end speech-to-text translation has been a great success, we argue that the cascaded speech-to-text translation model still has its pla...
Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has la...
Functional data - observations in the form of curves or trajectories - arise in diverse domains such as biomedical sensing, motion capture, and hand...
This paper presents Contourformer, a real-time contour-based instance segmentation algorithm. The method is fully based on the DETR paradigm and ach...
There has been substantial progress in humanoid robots, with new skills continuously being taught, ranging from navigation to manipulation. While th...
Robotic autonomy at centimeter scale requires compact and miniaturization-friendly actuation integrated with sensing and neural network processing a...
Single-view novel view synthesis (NVS) is a notorious problem due to its ill-posed nature, and often requires large, computationally expensive appro...
With the growing demand for oriented object detection (OOD), recent studies on point-supervised OOD have attracted significant interest. In this pap...
Identifying anatomical landmarks in 3D dental models is vital for orthodontic treatment, yet manual placement is complex and time-consuming. Althoug...
Virtual film production requires intricate decision-making processes, including scriptwriting, virtual cinematography, and precise actor positioning...
Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality, with lung metastases being the most common site of distant spread and...
Deep learning methods based on Convolutional Neural Networks (CNNs) have shown great potential to improve early and accurate diagnosis of Alzheimer'...
Cephalometric analysis is essential for the diagnosis and treatment planning of orthodontics. In lateral cephalograms, however, the manual detection...
The Text-to-Image (T2I) diffusion model is one of the most popular models in the world. However, serving diffusion models at the entire image level ...
The accelerated MRI reconstruction poses a challenging ill-posed inverse problem due to the significant undersampling in k-space. Deep neural networ...
Multi-modal Large Language Models (MLLMs) exhibit impressive capabilities in 2D tasks, yet encounter challenges in discerning the spatial positions,...
The exponential growth of short-video content has ignited a surge in the necessity for efficient, automated solutions to video editing, with challen...