Objective: To compare the accuracy, cost, and speed of a low-cost, non-generative structured-judgment model and four general-purpose large language models (LLMs) for extracting smoking status, pack-years, and quit date from clinical notes to support ...
Conversational systems based on large language models may support point-of-care evidence retrieval, but most evaluations use static benchmarks and few involve practicing physicians in low- and middle-income settings. We assessed the validity of answe...
Mammography interpretation requires tying each finding to supporting evidence and withholding judgment when that evidence is unavailable. Vision-language models are increasingly used for image-based medical question answering, yet existing benchmarks...
INTRODUCTION: Voice data combined with large language models may detect cognitive impairment, yet an interpretable framework has not been formally established. We developed and validated a scalable and interpretable framework, the Neurocognitive Spee...
Background: Pre-operative triage of adnexal masses with serum CA-125, HE4, and ultrasound (O-RADS) has limited accuracy, contributing to unnecessary surgery. We develop and validate a first of its kind, biomarker panel-free plasma proteomic classifie...
Aims. It is unknown whether large language models (LLMs) can predict a trial's result at the design stage, from the information available when it is registered. We tested three LLMs prospectively on the Late-Breaking Science programme of the European...
Background: Artificial intelligence (AI) is increasingly used in healthcare, but data on its real-world use and perceptions among obstetricians and gynecologists remain limited. Objectives: To assess familiarity, use, perceptions, and expectations re...
Background: General-purpose large language models (LLMs) are increasingly evaluated in diagnostic pathology, but prior studies have largely emphasized diagnostic accuracy rather than how models fail. We evaluated four LLMs for diagnostic performance,...
Background Artificial intelligence computer-aided diagnosis (AI CAD) algorithms for screening mammography have shown promise, but independent head-to-head comparisons of commercial algorithms on large, diverse cohorts remain limited. Methods 786,124 ...
BackgroundEarly-life air pollution exposure has been associated with childhood asthma, but its relative contribution is poorly characterised in cohorts with repeated indoor measurements. Short sampling windows and seasonal bias limit how well any sin...
Join thousands of healthcare professionals staying informed about the latest AI breakthroughs in medicine. Get curated insights delivered to your inbox.