Latest AI and machine learning research in patient safety / risk management for healthcare professionals.
Systematic reviews, scoping reviews, mapping studies, and related evidence syntheses are increasingly difficult to conduct with fully manual workflows as search volumes, update cycles, and synthesis requirements continue to expand. At the same time, artificial intelligence, machine learning, and large language models are rapidly entering review practice across query formulation, screening, extract...
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-base...
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both eviden...
Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics...
Background. Designing high-quality Objective Structured Clinical Examination (OSCE) stations is a time-consuming process. Generative artificial intell...
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the...
This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models. The recent advancements in multimodal models h...
Medical imaging is a cornerstone of diagnostics, yet automated chest X-ray report generation struggles with structural adherence, anatomical completen...
Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answ...
Existing evaluations of healthcare AI often treat interoperability as a technical infrastructure issue rather than a factor that directly influences t...
Progress in colonoscopy polyp segmentation is routinely reported through leaderboard comparisons on a small set of public benchmarks. We argue that th...
Objectives: To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest...
Background: Polypharmacy is common in people living with dementia (PLwD) and associated with adverse outcomes. Although Structured Medication Reviews ...
The growing demand for privacy-preserving data sharing has positioned synthetic data generation as a critical component of responsible AI workflows. D...
Language models increasingly write probabilistic programs (in NumPyro, Stan, or Pyro), but a program that compiles, runs, and passes every unit test c...
While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous sema...
Graph neural networks have moved from a niche representation-learning technique to the default model class wherever data carry relational structure. T...
Background: Guiding risk-appropriate inpatient thromboprophylaxis requires venous thromboembolism (VTE) risk stratification; however, reliable risk de...
Importance. Large language models (LLMs) increasingly inform mental health decisions by patients and clinicians. Inference-time activation steering ca...
Large Vision Language Models (VLMs) excel at visual question answering and semantic grounding, but their capacity for autonomous logical reasoning rem...