Latest AI and machine learning research in cultural competence for healthcare professionals.
Classifier-Free Guidance (CFG) is a widely used technique for improving conditional diffusion models by linearly combining the outputs of conditional and unconditional denoisers. While CFG enhances visual quality and improves alignment with prompts, it often reduces sample diversity, leading to a challenging trade-off between quality and diversity. To address this issue, we make two key contribu...
Recent advances in generative AI have enabled visual content creation through text-to-image (T2I) generation. However, despite their creative potential, T2I models often replicate and amplify societal stereotypes -- particularly those related to gender, race, and culture -- raising important ethical concerns. This paper proposes a theory-driven bias detection rubric and a Social Stereotype Index...
Temporal Logic (TL), especially Signal Temporal Logic (STL), enables precise formal specification, making it widely used in cyber-physical systems s...
We introduce Deep Spectral Prior (DSP), a new formulation of Deep Image Prior (DIP) that redefines image reconstruction as a frequency-domain alignm...
This paper examines how large language models (LLMs) are transforming core quantitative methods in communication research in particular, and in the ...
Recent advances in video generation models have sparked interest in world models capable of simulating realistic environments. While navigation has ...
We investigate the existence and persistence of a specific type of gender bias in some of the popular LLMs and contribute a new benchmark dataset, R...
Natural images exhibit label diversity (clean vs. noisy) in noisy-labeled image classification and prevalence diversity (abundant vs. sparse) in lon...
Recent advances in Multimodal Large Language Models (MLLMs) have shown promising results in integrating diverse modalities such as texts and images....
Person re-identification (ReID) models are known to suffer from camera bias, where learned representations cluster according to camera viewpoints ra...
As the marginal cost of scaling computation (data and parameters) during model pre-training continues to increase substantially, test-time scaling (...
Despite the impressive performance of generative Diffusion Models (DMs), their internal working is still not well understood, which is potentially p...
Large-scale Vision Language Models (LVLMs) are increasingly being applied to a wide range of real-world multimodal applications, involving complex v...
The biases exhibited by text-to-image (TTI) models are often treated as independent, though in reality, they may be deeply interrelated. Addressing ...
Segment Anything Models (SAM) have achieved remarkable success in object segmentation tasks across diverse datasets. However, these models are predo...
Medical Visual Question Answering (MedVQA) is crucial for enhancing the efficiency of clinical diagnosis by providing accurate and timely responses ...
BACKGROUND: The English National Health Service (NHS) strives for a fair, diverse, and inclusive workplace, but Black and Minority Ethnic (BME) repres...
Bias in Large Language Models (LLMs) significantly undermines their reliability and fairness. We focus on a common form of bias: when two reference ...
Human facial images encode a rich spectrum of information, encompassing both stable identity-related traits and mutable attributes such as pose, exp...
Although existing CLIP-based methods for detecting AI-generated images have achieved promising results, they are still limited by severe feature red...