Drug or Pokemon? An analysis of the ability of large language models to discern fabricated medications
Journal:
medRxiv
Published Date:
Apr 23, 2026
Abstract
Background Large language models (LLMs) are increasingly used in medication-related tasks despite limited evidence supporting their accuracy and safety. LLMs are vulnerable to adversarial attacks and may confabulate in response to fabricated information in the prompt. Given the substantial risk of harm from medication errors, the purpose of this study was to establish a benchmark task to evaluate LLM performance in distinguishing fabricated from real medications. Methods Two datasets (brand and generic), each consisting of 250 medication lists were developed, each including 4-6 medications and one fabricated medication (specifically a Pokemon character), with complete dosing information provided. Six LLMs were evaluated (GPT-5-Chat, GPT-4o-mini, DrugGPT, Gemma-3-27B-IT, Llama-3.3-70B-Instruct, and Qwen3-32B). For each medication list, models were queried for dosing information and disease indication across three experimental conditions: standard decoding, standard decoding with mitigation prompting, and deterministic decoding (temperature = 0). Each experiment was performed in triplicate. Outputs were deemed confabulations if the LLM failed to recognize and reject the fabricated medication. The primary outcome was the rate of confabulations. Exact paired-permutation tests were used to compare confabulation rates across LLMs and prompting approaches. Results Confabulation rates under standard conditions ranged from 2.7% to 99.6%. DrugGPT demonstrated the lowest baseline confabulation rates (2.7-6.4%), whereas other models demonstrated substantially higher rates (47.2-99.6%). Mitigation prompting significantly reduced confabulation rates compared to standard and deterministic decoding across all tasks and datasets (p < 0.001): GPT-5-Chat achieved confabulation rates of 0% with mitigation prompting from baseline rates of 66.4-88.5%. Deterministic decoding did not substantially impact confabulation rates. Conclusions LLMs performing medication-related tasks are susceptible to confabulation in response to invalid or fabricated inputs, raising important safety concerns for clinical use. While mitigation prompting reduced confabulation rates, this study underscores the importance of robust safeguards when deploying LLMs in medication-related decision-making.