Auditing Large Language Model-Generated Digital Standardized Patients for Demographic Bias: A Simulation Study with HIV Pre-Exposure Prophylaxis Screening as a Tracer Condition
Journal:
medRxiv
Published Date:
Sep 7, 2026
Abstract
Large language models (LLMs) are entering clinical training as digital standardized patients (DSPs), simulated patient encounters the model scripts and portrays. Demographic associations learned from corpus co-occurrences can enter at two points: the written case scripts and the live portrayals improvised in role-plays. We audited both pathways using an HIV pre-exposure prophylaxis (PrEP) screening scenario across six demographic factors (age, gender, marital status, sexual orientation, education, and race or ethnicity) in a 216-cell factorial design. With a generated arm and a template-substituted control arm, we simulated 4,320 conversations with fixed audit questions, measuring three channels: the case scripts, the composite role-play trainees receive, and role-play under identical control cases. Differences concentrated where corpus associations were strongest and reinforced predictable stereotypes. Anal sex appeared in 100% of gay mens cases, 83% of bisexual mens and 3% of heterosexual mens; the six factors explained 32% of the variance in composite sexual risk; and education predicted assigned socioeconomic status (adjusted R2 = 0.59). Demographics predicted 5 of the 9 improvised probe responses in composite role-play, and 2 under control cases. These two pathways sometimes had opposite signed effects: the model wrote women with higher alcohol use, but role-played them as drinking less, and the two canceled in composite role-play. An audit of either pathway alone would have obscured both effects. These issues are fixable as they are predictable and consistent. LLM-generated DSPs must have their cases and live portrayals audited as separate objects before deployment.