Latest AI and machine learning research in staffing & scheduling for healthcare professionals.
We present M2-omni, a cutting-edge, open-source omni-MLLM that achieves competitive performance to GPT-4o. M2-omni employs a unified multimodal sequence modeling framework, which empowers Large Language Models(LLMs) to acquire comprehensive cross-modal understanding and generation capabilities. Specifically, M2-omni can process arbitrary combinations of audio, video, image, and text modalities a...
Background: Recruitment for cohorts involving complex liver diseases, such as hepatocellular carcinoma and liver cirrhosis, often requires interpreting semantically complex criteria. Traditional manual screening methods are time-consuming and prone to errors. While AI-powered pre-screening offers potential solutions, challenges remain regarding accuracy, efficiency, and data privacy. Methods: We...
Key-value (KV) caching has emerged as a crucial optimization technique for accelerating inference in large language models (LLMs). By allowing the a...
Domain shift presents a significant challenge in applying Deep Learning to the segmentation of 3D medical images from sources like Magnetic Resonanc...
Advances in Large Language Models revolutionized medical education by enabling scalable and efficient learning solutions. This paper presents a pipe...
Diffusion models (DMs) have revolutionized data generation, particularly in text-to-image (T2I) synthesis. However, the widespread use of personaliz...
We present mean-shift distillation, a novel diffusion distillation technique that provides a provably good proxy for the gradient of the diffusion o...
The application of large language models (LLMs) in healthcare has the potential to revolutionize clinical decision-making, medical research, and pat...
Motor skill acquisition in fields like surgery, robotics, and sports involves learning complex task sequences through extensive training. Traditiona...
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a range of multimodal tasks. However, their inference efficien...
Large language models (LLMs) have been widely adopted in various downstream task domains. However, their ability to directly recall and apply factua...
Understanding the relationship between the evolution of microstructures of irradiated LiAlO2 pellets and tritium diffusion, retention and release co...
Due to the widespread use of LLMs and the rising critical ethical and safety concerns, LLM unlearning methods have been developed to remove harmful ...
Mentoring software is a pivotal innovation in addressing critical challenges in teacher development within educational institutions. This study expl...
Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple s...
Transformers have become the backbone of neural network architecture for most machine learning applications. Their widespread use has resulted in mu...
Automating planning with LLMs presents transformative opportunities for traditional industries, yet remains underexplored. In commercial constructio...
Large language models (LLMs) require immense resources for training and inference. Quantization, a technique that reduces the precision of model par...
We present SkyReels-A1, a simple yet effective framework built upon video diffusion Transformer to facilitate portrait image animation. Existing met...
Facial expression recognition (FER) systems in low-resolution settings face significant challenges in accurately identifying expressions due to the ...