Latest AI and machine learning research in cultural competence for healthcare professionals.
Bias and inequity in palliative care disproportionately affect marginalised groups. Large language models (LLMs), such as GPT-4o, hold potential to enhance care but risk perpetuating biases present in their training data. This study aimed to systematically evaluate whether GPT-4o propagates biases in palliative care responses using adversarially designed datasets. In July 2024, GPT-4o was probed...
While recent work has found that vision-language models trained under the Contrastive Language Image Pre-training (CLIP) framework contain intrinsic social biases, the extent to which different upstream pre-training features of the framework relate to these biases, and hence how intrinsic bias and downstream performance are connected has been unclear. In this work, we present the largest compreh...
Text-to-image generation models have gained popularity among users around the world. However, many of these models exhibit a strong bias toward Engl...
With the rapid development of AI-generated content (AIGC), the creation of high-quality AI-generated videos has become faster and easier, resulting ...
Customized generation has achieved significant progress in image synthesis, yet personalized video generation remains challenging due to temporal in...
Hairstyles are intricate and culturally significant with various geometries, textures, and structures. Existing text or image-guided generation meth...
Recent advancements in reinforcement learning (RL) have achieved great success in fine-tuning diffusion-based generative models. However, fine-tunin...
Can we derive computational metrics to quantify visual creativity in drawings across intelligent agents, while accounting for inherent differences i...
This paper provides guidance for building and maintaining infrastructure for participatory AI efforts by sharing reflections on building World Wide ...
The generation of incorrect images, such as depictions of people of color in Nazi-era uniforms by Gemini, frustrated users and harmed Google's reput...
While diffusion models are powerful in generating high-quality, diverse synthetic data for object-centric tasks, existing methods struggle with scen...
Current Vision Language Models (VLMs) remain vulnerable to malicious prompts that induce harmful outputs. Existing safety benchmarks for VLMs primar...
Image generation abilities of text-to-image diffusion models have significantly advanced, yielding highly photo-realistic images from descriptive te...
India's rich cultural and linguistic diversity poses various challenges in the domain of Natural Language Processing (NLP), particularly in Named En...
Single-domain generalization for object detection (S-DGOD) aims to transfer knowledge from a single source domain to unseen target domains. In recen...
The proliferation of Text-to-Image (T2I) models has revolutionized content creation, providing powerful tools for diverse applications ranging from ...
Unified multimodal large language models (U-MLLMs) have demonstrated impressive performance in visual understanding and generation in an end-to-end ...
This paper reports on the results from a pilot study investigating the impact of automatic speech recognition (ASR) technology on interpreting quali...
Text-embedding models often exhibit biases arising from the data on which they are trained. In this paper, we examine a hitherto unexplored bias in ...
Survival analysis, a vital tool for predicting the time to event, has been used in many domains such as healthcare, criminal justice, and finance. L...