State Required CME

Patient safety / Risk Management

Latest AI and machine learning research in patient safety / risk management for healthcare professionals.

10,112 articles
Stay Ahead - Weekly Patient safety / Risk Management research updates
Subscribe
Browse Categories
Showing 561-580 of 10,112 articles

ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies

Systematic reviews, scoping reviews, mapping studies, and related evidence syntheses are increasingly difficult to conduct with fully manual workflows as search volumes, update cycles, and synthesis requirements continue to expand. At the same time, artificial intelligence, machine learning, and large language models are rapidly entering review practice across query formulation, screening, extract...

Aug 25 2026 2608.25050v1

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-base...

Aug 21 2026 2608.20886v1
G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both eviden...

Aug 20 2026 2608.20331v1
CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering

Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics...

Aug 11 2026 2608.11074v1
Characterizing large language model generative artificial intelligence variability in the production of objective structured clinical examination stations

Background. Designing high-quality Objective Structured Clinical Examination (OSCE) stations is a time-consuming process. Generative artificial intell...

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the...

Aug 3 2026 2608.02589v1
Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models

This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models. The recent advancements in multimodal models h...

Jul 24 2026 2607.21998v1
DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization

Medical imaging is a cornerstone of diagnostics, yet automated chest X-ray report generation struggles with structural adherence, anatomical completen...

Jul 21 2026 2607.18988v1
Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks

Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answ...

Jul 14 2026 2607.12278v1
FHIRTrustBench: A Benchmark for Interoperability-Driven Clinical AI Readiness and Trustworthiness

Existing evaluations of healthcare AI often treat interoperability as a technical infrastructure issue rather than a factor that directly influences t...

Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks

Progress in colonoscopy polyp segmentation is routinely reported through leaderboard comparisons on a small set of public benchmarks. We argue that th...

Jul 9 2026 2607.08203v1
Diagnostic accuracy of a DenseNet-121 deep learning algorithm for chest radiograph triage in health assessment applicants: a prospective shadow-mode validation study in Nepal

Objectives: To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest...

Attitudes of People Living with Dementia and their Carers towards the use of Generative Artificial Intelligence to inform Structured Medication Reviews

Background: Polypharmacy is common in people living with dementia (PLwD) and associated with adverse outcomes. Although Structured Medication Reviews ...

TDGT: A Tabular Data Generation Toolkit supporting adaptive GPU-accelerated Bayesian mixture models, diffusion-based models, and latent-space generative modeling

The growing demand for privacy-preserving data sharing has positioned synthetic data generation as a critical component of responsible AI workflows. D...

Jun 30 2026 2606.31268v1
Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

Language models increasingly write probabilistic programs (in NumPyro, Stan, or Pyro), but a program that compiles, runs, and passes every unit test c...

Jun 30 2026 2606.31630v1
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous sema...

Jun 29 2026 2606.30124v1
Graph Neural Networks Applications Across Domains: All Insights You Need

Graph neural networks have moved from a niche representation-learning technique to the default model class wherever data carry relational structure. T...

Jun 25 2026 2606.27202v1
Multisite Real-World Validation of an Electronic Health Record-Integrated Generative Artificial Intelligence Tool for Venous Thromboembolism Risk Stratification

Background: Guiding risk-appropriate inpatient thromboprophylaxis requires venous thromboembolism (VTE) risk stratification; however, reliable risk de...

Silent Manipulation of Mental Health Treatment Recommendations from a Large Language Model

Importance. Large language models (LLMs) increasingly inform mental health decisions by patients and clinicians. Inference-time activation steering ca...

LADBench: A Benchmark for Logical Fault Detection in Images

Large Vision Language Models (VLMs) excel at visual question answering and semantic grounding, but their capacity for autonomous logical reasoning rem...

Jun 16 2026 2606.17433v1
Browse Categories