On the robustness of medical term representations in locally deployable language models
Journal:
medRxiv
Published Date:
Feb 26, 2026
Abstract
Background: Hosting large language models (LLMs) on premises can secure patient data but requires compact architectures to function on standard hardware. The impact of such constraints on the robustness of their representations for medical terminology is important for clinical AI safety but poorly understood. The statistical nature of LLM training inherently limits the representation of terms with low societal prominence or lexical frequency, and high ambiguity. Methods: We assessed 15 open weights LLMs (4B to 120B) for their representational robustness of 250 neurological terms. Neurology was chosen for its strict hierarchical and anatomical terminology. A term representation was deemed robust only if the model correctly navigated four tests, verifying valid links against distractors and reverse associations. We examined associations between representational robustness and model size, medical fine tuning, and five terminological subdomains (localisation, clinical features, investigations, diagnoses, and treatments). We assessed term difficulty using the semantic complexity index (SCI), a novel composite integrating societal prominence, lexical frequency, and ambiguity. Results: Representational robustness followed a log linear scaling law relative to model size (r=0.736, p=0.002). Medical finetuning yielded no benefit for 4B models, but significantly improved larger 27B model performance, with rate of robust representations rising from 38.2% to 62.6% (p<0.0001). While most local LLMs performance degraded sharply with increasing SCI values, GPT OSS 20B and 120B maintained complexity invariance (with <20% decline from lowest to highest complexity terms). Notably, the general purpose 20B GPT OSS model outperformed larger and medically fine tuned counterparts. Robustness varied by subdomain (F=4.69, p=0.003), with diagnoses (73.8%) scoring significantly higher than localisation (47.9%, p=0.004) and clinical features (52.1%, p=0.02). Conclusions: While representational robustness broadly follows model size scaling laws, neither model size nor finetuning guarantees clinical reliability. Since performance fluctuates with terminological complexity and subdomain, safe deployment requires validating representational robustness for specific use cases rather than assuming larger models handle medical language safely.