MTS-Bench: A Manchester Triage System Benchmark for Language Model Triage Safety
Journal:
medRxiv
Published Date:
Aug 5, 2026
Abstract
Background: General purpose language models such as ChatGPT are increasingly used by physicians and triage nurses during emergency triage. A recent study reported 51.6% undertriage of emergencies when patients queried ChatGPT directly (Ramaswamy et al., 2026). DR. INFO is an agentic AI based clinical assistant that retrieves over a curated clinical knowledge base, and an MTS specific retrieval configuration is available in which the system also retrieves the Manchester Triage System (MTS) textbook at inference time. The safety of these systems as a triage adjunct against a structured framework has not been characterised. Methods: We adapted the clinical scenarios published by Ramaswamy et al. and mapped them to the Manchester Triage System, yielding 39 emergency cases covering all five MTS priority levels. Each case was evaluated in two variants, one without and one with the objective clinical data block (vital signs, examination findings, and laboratory results), and permuted across two genders, giving 156 prompts per condition. Three systems were tested with and without a misleading GP referral statement prepended as an anchoring statement, giving 312 prompts per system: DR. INFO Baseline, DR. INFO with MTS retrieval, and OpenAI GPT-5.1. The primary outcome was the undertriage rate on the ordered MTS scale, tested with Fisher's exact test. Results: GPT-5.1 undertriaged 44.2% of cases (69/156; 95% CI 36.7 to 52.1), including 75.0% of Red and 73.4% of Orange presentations. Both DR. INFO configurations undertriaged 11.5% of cases (18/156; 95% CI 7.4 to 17.5; Fisher's exact p = 1.0 x 10^-10 versus GPT-5.1). GPT-5.1 produced 6 dangerous misses (3.8%), and both DR. INFO configurations produced none (p = 0.030). When the anchoring statement was prepended, GPT-5.1 undertriaged 8 of 8 Red cases, while both DR. INFO configurations continued to undertriage none. Adding objective clinical data to the input reduced undertriage in DR. INFO with MTS retrieval from 19.2% to 3.8% (p = 0.005). DR. INFO Baseline and GPT-5.1 showed no comparable change. There was no significant effect of gender. Conclusion: On this benchmark, replacing a general purpose language model with an agentic retrieval augmented system over a curated clinical knowledge base substantially reduced the undertriage and dangerous miss rates. Adding retrieval of the Manchester Triage System textbook to the agentic system was further associated with a reduced susceptibility to the anchoring statement and with an appropriate change in the assigned MTS priority when objective clinical data became available. Of the three configurations evaluated here, only DR. INFO with MTS retrieval combined a clinically conservative assignment at first contact with appropriate updating as additional clinical information arrived.