AIMC Journal:
medRxiv

Showing 11 to 20 of 4068 articles

Validation of a two-stage automated screening pipeline for medical systematic reviews: a stratified concordance study with three independent human reviewers, using PABAK and Gwet's AC1

medRxiv
Background. Large language models (LLMs) are increasingly used to automate title-abstract screening in systematic reviews, but validation studies typically report sensitivity, specificity, or Cohen's kappa: metrics that are unstable under the low and...

HealthFound: a health world model for quantitative reasoning on longitudinal health profiles

medRxiv
Medical large language models (LLMs) have shown promise in medical knowledge retrieval and clinical reasoning. However, their capacity for quantitative reasoning over dense, longitudinal health data remains limited, largely due to the lack of effecti...

RFM: A Lightweight Retinal Foundation Model for Generalised Oculomics

medRxiv
Foundation models for colour fundus photography (CFP) are candidate visual encoders for multimodal medical AI, but existing models are large and are assessed under fine-tuning rather than the frozen-encoder conditions such systems impose. We present ...

Forecasting Alzheimer's disease progression using non-imaging clinical features and XGBoost

medRxiv
Identifying progressive dementia during its prodromal stage is a challenging yet critical task for healthcare providers. Traditional clinical assessments are often not sensitive enough to detect warning signs of cognitive decline. Machine learning (M...

Generalizing Amyloid β Pathology from Postmortem Human Brains with Vision Foundation Models for Lesion Quantification

medRxiv
Background: Traditional neuropathological assessment of Alzheimer's disease (AD) relies on manual, semi-quantitative scoring of amyloid-beta (A{beta}) burden. This approach is limited by high inter-rater variability and often qualitative, failing to ...

Multimodal Artificial Intelligence Predicts Pathological Response and Prognosis in HER2-Positive Early Breast Cancer Treated with Chemotherapy De-escalation Strategies: Analyses from PHERGain and PHERGain-2 Trials

medRxiv
Background Chemotherapy (CT) de-escalation strategies based on dual HER2 blockade with trastuzumab and pertuzumab (HP) have shown promising efficacy in HER2-positive (HER2+) early breast cancer (EBC), but biomarkers to guide patient selection are lac...

Subversion of clinical judgment by conversational artificial intelligence

medRxiv
In medicine, artificial intelligence (AI) safety has focused on discrete flaws such as hallucination, bias and miscalibration. Conversational systems pose a distinct hazard: in dialogue, a model can covertly advance a misaligned objective by redirect...

Can Generative Video Address the Clinical Video Data Gap? Evaluation of Synthetic Parkinsonian Hand Motions

medRxiv
Computer vision approaches to disease motor assessment are limited by clinical video dataset scarcity and distributional imbalance, motivating interest in generative video models as a potential source of training data. We introduce a three-component ...

Machine Learning Models Using Electronic Health Record Data to Predict Obstructive Sleep Apnea

medRxiv
Rationale: Obstructive sleep apnea (OSA) is highly prevalent yet largely under-diagnosed. Current screening strategies rely on effortful questionnaires with modest accuracy, and prior machine learning models often require resource-intensive inputs an...

Feature-Engineering Strategies for EEG-Based ADHD Classification in Children: A Controlled Benchmark of Expert, LLM-Guided and Full-Space Comparators

medRxiv
EEG-based machine-learning studies of attention-deficit/hyperactivity disorder (ADHD) are vulnerable to optimistic performance estimates when feature selection or model choice is informed by data outside the training fold. We compared expert-reconstr...