AIMC Journal:
medRxiv

Showing 101 to 110 of 4073 articles

Extracting smoking history from clinical notes for lung cancer screening decision support: comparing a structured-judgment model with general-purpose large language models

medRxiv
Objective: To compare the accuracy, cost, and speed of a low-cost, non-generative structured-judgment model and four general-purpose large language models (LLMs) for extracting smoking status, pack-years, and quit date from clinical notes to support ...

Quality of physicians responses using a conversational artificial intelligence system: a randomized vignette experiment in Latin America

medRxiv
Conversational systems based on large language models may support point-of-care evidence retrieval, but most evaluations use static benchmarks and few involve practicing physicians in low- and middle-income settings. We assessed the validity of answe...

CARE-MVLM: Counterfactual Abstention and Region-Grounded Evidence in Mammography Vision-Language Models

medRxiv
Mammography interpretation requires tying each finding to supporting evidence and withholding judgment when that evidence is unavailable. Vision-language models are increasingly used for image-based medical question answering, yet existing benchmarks...

A neurocognitive speech taxonomy for voice biomarkers of Alzheimer's disease

medRxiv
INTRODUCTION: Voice data combined with large language models may detect cognitive impairment, yet an interpretable framework has not been formally established. We developed and validated a scalable and interpretable framework, the Neurocognitive Spee...

Differentiating benign from malignant adnexal masses by biomarker-agnostic plasma proteomics using adaptive machine learning

medRxiv
Background: Pre-operative triage of adnexal masses with serum CA-125, HE4, and ultrasound (O-RADS) has limited accuracy, contributing to unnecessary surgery. We develop and validate a first of its kind, biomarker panel-free plasma proteomic classifie...

A pre-registered prospective study of large language models predicting late-breaking cardiovascular trial results at ESC Congress 2026

medRxiv
Aims. It is unknown whether large language models (LLMs) can predict a trial's result at the design stage, from the information available when it is registered. We tested three LLMs prospectively on the Late-Breaking Science programme of the European...

Use and Perceptions of Artificial Intelligence Among Respondents to a Self-Selected Survey of Obstetrician-Gynecologists in France

medRxiv
Background: Artificial intelligence (AI) is increasingly used in healthcare, but data on its real-world use and perceptions among obstetricians and gynecologists remain limited. Objectives: To assess familiarity, use, perceptions, and expectations re...

Errors, Hallucinations, and Clinical Impact of General-Purpose Multimodal Large Language Models in Histopathology

medRxiv
Background: General-purpose large language models (LLMs) are increasingly evaluated in diagnostic pathology, but prior studies have largely emphasized diagnostic accuracy rather than how models fail. We evaluated four LLMs for diagnostic performance,...

Kaiser Permanente National Cross-Vendor Validation of Mammography Artificial Intelligence Computer-Aided Diagnosis Algorithms in a US-Representative Population

medRxiv
Background Artificial intelligence computer-aided diagnosis (AI CAD) algorithms for screening mammography have shown promise, but independent head-to-head comparisons of commercial algorithms on large, diverse cohorts remain limited. Methods 786,124 ...

Associations Between Early-Life Indoor and Outdoor Air Pollution Exposure and Childhood Asthma

medRxiv
BackgroundEarly-life air pollution exposure has been associated with childhood asthma, but its relative contribution is poorly characterised in cohorts with repeated indoor measurements. Short sampling windows and seasonal bias limit how well any sin...