TCM LLM · Datasets
TCM LLM datasets
Public datasets for training and evaluating TCM LLMs: instruction data, case QA, formulas, classical corpora, and benchmarks.
TCM-RobustSDT
TCM-RobustSDT: a robustness benchmark dataset for LLM clinical reasoning in TCM (Figshare).
HSQ-TD(健身气功指令微调数据集)
HSQ-TD: the first instruction-tuning dataset for health Qigong/wellness, with 57,843 instructions distilled from official textbooks and professional literature (ScienceDB).
TCM_KG
ChatMed knowledge graph.
TCM-MKG
TCM multi-dimensional knowledge graph.
OpenTCM-KG
OpenTCM gynecology classics KG (~48k entities / ~152k relations).
TCMNSCLC
Real-world NSCLC TCM reasoning dataset with fully annotated cases (pattern differentiation / treatment method / decoction / patent medicine).
ChP-TCM
KnowledgeQA and PrescriptionWriting instructions built from Chinese Pharmacopoeia Vol. I.
Traditional-Chinese-Medicine-Dataset-SFT
High-quality TCM supervised fine-tuning dataset.
TCMChat-dataset-600k
TCMChat herbal QA and recommendation instruction data (~600k).
TCM-Instruction-Tuning-ShizhenGPT
ShizhenGPT multimodal SFT data (text/vision/speech/ECG etc.; ~311k items total per paper Table 3).
ShenNong_TCM_Dataset
ShenNong TCM instruction dataset.
MedChatZH
MedChatZH TCM consultation dataset.
ChatMed_Consult_Dataset
Chinese online medical consult dataset (500k+ consults with ChatGPT replies).
CMtMedQA
ZhongJing real multi-turn doctor–patient dialogues (~70k).
Baize-TCM-Corpus-V3
~157k TCM QA items covering theory, herbs, formulas, diagnosis, acupuncture, and clinic.
neijing-sft-v1.2
~2,009 Neijing-related instruction samples for Xinghe, with thinking/output fields.
TCM-Text-Exams
Recent TCM licensure / graduate-exam text benchmark.
Medical-LLMs-Chinese-Exam
Chinese medical exam evaluation for medical LLMs.
ZhongJing-OMNI
ZhongJing-OMNI multimodal TCM eval (including tongue).
TCMEval-SDT
TCMEval-SDT: a benchmark of 300 syndrome-diagnosis cases (web, classical texts, hospital records) for evaluating TCM syndrome-differentiation reasoning, with FAIR metadata (Sci. Data 2025).
TCMBench
TCMBench: a comprehensive benchmark for evaluating LLMs in traditional Chinese medicine (arXiv 2024).
TCM-Vision-Benchmark
TCM vision benchmark (herb recognition / inspection, ~7k items).
TCM-Tongue
6,719 standardized tongue images with 20-class multi-label pathology annotations and detection baselines.
TCM-Ladder
TCM-Ladder: a multimodal QA benchmark for comprehensively evaluating TCM multimodal LLMs on real-world tasks (arXiv 2025).
TCM-Eval
Dynamic, extensible TCM evaluation platform.
TCM-BEST4SDT
Case benchmark for syndrome differentiation and treatment.
TCM-5CEval
Five-dimension deep TCM evaluation suite.
TCM-3CEval
Three-axis eval: core knowledge, classics, clinical decisions.
MTCMB
MTCMB dataset: a multi-task TCM benchmark covering knowledge, reasoning and safety, 12 subsets with ~7,100 samples (arXiv 2025).
HWTCMBench
HWTCMBench TCM capability evaluation set.
TCMEval-PA
328 multiple-choice items on prescription normative quality and safety auditing.
LingLan
LingLan large multi-task TCM evaluation benchmark (2026).
ChiMed 2.0
Upgraded Chinese medical pretraining dataset covering TCM corpora for LLM pretraining.
classical-tcm-canon
Full-text digitizations of the TCM canon: Neijing, Nanjing, Shanghan Lun, Jingui Yaolue and warm-disease classics.
Traditional-Chinese-Medicine-Dataset-Pretrain
High-quality TCM pretraining dataset from non-Internet sources (~1GB; clinical cases, classics, encyclopedia), 99% simplified Chinese.
TCM-Pretrain-Data-ShizhenGPT
ShizhenGPT pretraining corpus (15B+ tokens reported in the paper — Stage-1 text 11.92B incl. 6.3B TCM, plus Stage-2 multimodal ~3.6B).
TCM-Ancient-Books
A corpus of nearly 700 TCM ancient-book texts.
awesome_Chinese_medical_NLP
Curated list of Chinese medical NLP resources: terminologies, corpora, word vectors, pretrained models, KGs, NER and QA (incl. CBLUE).
CPM中成药数据集
Living large-scale public Chinese patent medicine data accompanying RAG-CPMF.