AIMC Journal:
medRxiv

Showing 331 to 340 of 4073 articles

Retrieval-Augmented Large Language Models for Clinically Aligned Adverse Event Coding in Acute Myeloid Leukemia Clinical Trials

medRxiv
Background: Adverse event (AE) coding is essential for safety monitoring in oncology clinical trials, particularly in acute myeloid leukemia (AML), where intensive therapies are associated with frequent and heterogeneous toxicities requiring standard...

Prediction of Heart Failure based on Multimodal Data from MIMIC-IV

medRxiv
Heart failure (HF) affects over 64 million people worldwide and remains a leading cause of cardiovascular mortality. Early identification of patients at risk is essential for timely treatment and to support hospital and primary care physicians. This ...

Robustness Gap of Large Language Models in Nephrology

medRxiv
Background: Whether benchmark performance reflects robust clinical reasoning rather than surface-level pattern recognition remains uncertain. We evaluated the robustness of state-of-the-art large language models (LLMs) on nephrology board renewal que...

Preoperative Prediction of Residual Cancer Burden After Neoadjuvant Chemotherapy in Breast Cancer: A Multimodal Machine Learning Approach and Implications for Clinical Decision Support

medRxiv
Background. Residual cancer burden (RCB) after neoadjuvant chemotherapy (NAC) offers finer prognostic stratification than binary pathologic complete response, and increasingly guides adjuvant treatment intensity. Predicting four-tier RCB class from p...

Community Learning Ledgers for Cancer Navigation in Small Island Developing States

medRxiv
Importance. Cancer is the second leading cause of death among patients in the Caribbean, where outcomes are associated with delayed clinical navigation to screening, diagnosis, and treatment. Artificial intelligence is increasingly used to guide pati...

From Output Errors to Workflow Harm: A Practitioner-Audit Method for LLM-Mediated Research

medRxiv
Objective. Formal large language model (LLM) evaluations score isolated prompts, but clinicians and health-informatics researchers meet model failures inside multi-step workflows where erroneous output can alter procedures or contaminate documents. W...

REFINE: Closing the Loop Between Large Language Models and Symbolic Rules in Clinical NLP

medRxiv
Symbolic clinical natural language processing (NLP) systems remain widely used for extracting clinical concepts from electronic health record (EHR) narratives, but maintaining rule resources requires extensive manual error analysis and rule refinemen...

Development and Internal Validation of a Large Language Model Pipeline for Multi-Label Classification of Patient Portal Messages

medRxiv
Objectives: Characterizing patient portal message content at scale can help target efforts to manage administrative work. We developed and validated a large language model (LLM) pipeline for multi-label classification of messages using an expert-deri...

Are automated documentation-error judges fit to measure ambient AI scribes? A pre-registered, blinded human-validation study

medRxiv
Objectives Safety claims for ambient artificial intelligence (AI) scribes rest on automated judges that detect documentation errors and grade clinical risk. Expert reviewers are under-sensitive and disagree with one another, so no gold standard exist...

Performance of an Ambient Generative AI Documentation Tool in a Linguistically Diverse Clinical Setting

medRxiv
Ambient artificial intelligence scribes are being increasingly used in healthcare to improve efficiency and reduce provider clinical documentation burden, yet their performance across linguistically diverse patient populations is not well characteriz...