Latest AI and machine learning research in cultural competence for healthcare professionals.
Large Multimodal Models (LMMs) have emerged as powerful models capable of understanding various data modalities, including text, images, and videos. LMMs encode both text and visual data into tokens that are then combined and processed by an integrated Large Language Model (LLM). Including visual tokens substantially increases the total token count, often by thousands. The increased input length...
Image captioning tasks usually use two-stage training to complete model optimization. The first stage uses cross-entropy as the loss function for optimization, and the second stage uses self-critical sequence training (SCST) for reinforcement learning optimization. However, the SCST algorithm has certain defects. SCST relies only on a single greedy decoding result as a baseline. If the model its...
Bayesian Neural Networks (BNNs) provide a promising framework for modeling predictive uncertainty and enhancing out-of-distribution robustness (OOD)...
Chronic diseases such as obesity and hypertension due to malnutrition can be prevented by following the appropriate diet, correct diet intake with cor...
INTRODUCTION: Artificial intelligence (AI) has significant potential to improve health outcomes in oncology. However, as AI utility increases, it is i...
Autoregressive (AR) modeling, known for its next-token prediction paradigm, underpins state-of-the-art language and visual generative models. Tradit...
Leveraging the effective visual-text alignment and static generalizability from CLIP, recent video learners adopt CLIP initialization with further r...
Face recognition (FR) stands as one of the most crucial applications in computer vision. The accuracy of FR models has significantly improved in rec...
Generative AI (e.g., ChatGPT) is increasingly integrated into people's daily lives. While it is known that AI perpetuates biases against marginalize...
Machine learning systems trained on electronic health records (EHRs) increasingly guide treatment decisions, but their reliability depends on the cr...
STEM fields are traditionally male-dominated, with gender biases shaping perceptions of job accessibility. This study analyzed gender representation...
Text-to-image diffusion models often exhibit biases toward specific demographic groups, such as generating more males than females when prompted to ...
The debate around bias in AI systems is central to discussions on algorithmic fairness. However, the term bias often lacks a clear definition, despi...
Large Language Models have garnered significant attention for their capabilities in multilingual natural language processing, while studies on risks...
This study addresses the challenge of reconstructing unseen ECG signals from PPG signals, a critical task for non-invasive cardiac monitoring. While...
Linker generation is critical in drug discovery applications such as lead optimization and PROTAC design, where molecular fragments are assembled in...
Data diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection...
Ensuring equitable Artificial Intelligence (AI) in healthcare demands systems that make unbiased decisions across all demographic groups, bridging t...
Federated learning (FL) has shown great potential in medical image computing since it provides a decentralized learning paradigm that allows multipl...
Traditional evaluations of multimodal large language models (LLMs) have been limited by their focus on single-image reasoning, failing to assess cru...