LLM-SDaT: A knowledge-informed LLM framework for syndrome differentiation in TCM.
Journal:
Neural networks : the official journal of the International Neural Network Society
Published Date:
Mar 4, 2026
Abstract
The rapid advancement of large language models (LLMs) offers significant potential for enhancing artificial intelligence applications in Traditional Chinese Medicine (TCM), particularly in improving the precision and interpretability of syndrome differentiation and treatment planning. However, current methods are often limited by the lack of standardized, large-scale clinical datasets and insufficient integration of structured TCM knowledge, which hinders diagnostic accuracy and generalizability. To address these issues, we propose LLM-SDaT, a knowledge-informed adaptation framework that incorporates domain-specific expertise into LLMs via parameter-efficient fine-tuning. Specifically, we first introduce two structured datasets: TCMSD100, a large-scale clinical corpus with annotated patient records across 100 syndromes, and TCMSDaT100, a knowledge-informed dataset integrating syndrome definitions, etiologies, and canonical prescriptions. Then, we implement a two-stage fine-tuning framework based on LoRA. The first stage, trained on TCMSD100, focuses on accurate syndrome differentiation. The second stage then aligns the model with the comprehensive knowledge in TCMSDaT100 to generate clinically coherent and personalized treatment recommendations. Experiments demonstrate the superiority of our approach, which achieves an F1-score of 85.19% in syndrome classification, significantly outperforming existing baselines and general-purpose LLMs. This work underscores the value of integrating structured knowledge through parameter-efficient adaptation, offering a scalable pathway toward interpretable TCM decision-support systems. All datasets and codes are publicly available at: TobyChain/TCMSDaT.git.
Authors
Keywords
No keywords available for this article.