Latest AI and machine learning research in cultural competence for healthcare professionals.
Cell state diversity drives tissue adaptability, repair, and disease resilience, but fully capturing this cellular complexity remains a challenge. Most current approaches rely on transcriptional profiling and often overlook functional insights embedded in organelle structure, key indicators of cellular metabolism and stress. Here, we introduce spatial Organellomics (sOrganellomics), an imaging wor...
Ship detection for navigation is a fundamental perception task in intelligent waterway transportation systems. However, existing public ship detection datasets remain limited in terms of scale, the proportion of small-object instances, and scene diversity, which hinders the systematic evaluation and generalization study of detection algorithms in complex maritime environments. To this end, we cons...
Referring Expression Segmentation (RES) aims to segment image regions described by natural-language expressions, serving as a bridge between vision an...
Ensuring fairness in machine learning predictions is a critical challenge, especially when models are deployed in sensitive domains such as credit sco...
RL training of multi-turn LLM agents is inherently unstable, and reasoning quality directly determines task performance. Entropy is widely used to tra...
Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growin...
Data heterogeneity hinders clinical deployment of medical image analysis models, and generative data augmentation helps mitigate this issue. However, ...
Generating realistic single-cell transcriptomic profiles from structured biological descriptions would enable controlled simulation, data augmentation...
Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, con...
Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. W...
The rapid proliferation of AI-generated images, powered by generative adversarial networks (GANs), diffusion models, and other synthesis techniques, h...
Pediatric bipolar disorder is challenging to diagnose accurately due to symptom heterogeneity. More standardized and data-driven approaches are needed...
Background: Large language models (LLMs) have been evaluated as tools to assist rare disease diagnosis, yet evidence on their accuracy remains fragmen...
We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasonin...
Multimodal Large Language Models (MLLMs) have recently been explored as face verification systems that determine whether two face images are of the sa...
Recent multimodal large language models are computationally expensive because Transformers must process a large number of visual tokens. We present \t...
Recent research has shown that text-to-image diffusion models are capable of generating high-quality images guided by text prompts. But can they be us...
Long-tail distributions in driving datasets pose a fundamental challenge for 3D perception, as rare classes exhibit substantial intra-class diversity ...
The expansion of generative AI and LLM services underscores the growing need for adaptive mechanisms to select an appropriate available model to respo...
Autoregressive (AR) models are highly effective for image generation, yet their standard maximum-likelihood estimation training lacks direct optimizat...