TCM LLM · Datasets
TCM LLM datasets
Public datasets for training and evaluating TCM LLMs: instruction data, case QA, formulas, classical corpora, and benchmarks.
TCM-RobustSDT
TCM-RobustSDT: a robustness benchmark dataset for LLM clinical reasoning in TCM (Figshare).
HSQ-TD(健身气功指令微调数据集)
HSQ-TD: the first instruction-tuning dataset for health Qigong/wellness, with 57,843 instructions distilled from official textbooks and professional literature (ScienceDB).
OASIS (KIOM)
KIOM traditional-medicine literature portal for Korean-medicine papers and herbal resources.
Korean Medicine Embedding Dataset
Query–positive–negatives (~113k pairs) built from Korean-medicine terms and an ontology, for embedding fine-tunes such as BGE-M3.
KNApSAcK KAMPO
NAIST Kampo public database (~1,581 formulas, 278 crude drugs), downloadable from the NBDC life-science archive.
webMedQA
Early Chinese non-factoid medical QA from health-consult sites (~63k questions, one positive and four negative answers each).
IMCS-21
About 4,116 pediatric online consults annotated for entities, intents, symptoms, and reports; later wired into four CBLUE dialogue tasks.
CBLUE
Chinese biomedical NLU benchmark (NER, relations, diagnosis normalization, classification); the source-task suite behind PromptCBLUE, with a Tianchi submission portal.
CMeKG
Chinese medical knowledge graph of diseases, drugs, and symptoms; main source for BenCao/HuaTuo and ChatGLM-Med instruction data. Official portal is unstable; verify via the tools repo.
MedDialog
Large doctor–patient dialogue corpus (about 1.1M Chinese encounters) used for multi-turn Chinese medical fine-tuning.
cMedQA2
Chinese community medical QA (~108k questions / 200k answers), a common source for BianQue-style SFT mixtures.
cMedQA
Chinese community medical QA-matching set (repo table ~54k questions / 102k answers; non-commercial research). Paper DOI matches the README; see cMedQA2 for the later release.
ChiMed (Qilin)
Qilin-Med's ~3GB Chinese medical corpus (CPT/SFT/DPO); not the same resource as the ChiMed 2.0 pretraining set.
DISC-Med-SFT
Fudan DISC medical-dialogue SFT set (~470k examples from KG triples and reconstructed consults; no preference data).
PromptCBLUE
CBLUE's 16 Chinese medical NLP tasks rewritten as generative instructions; an early unified Chinese medical LLM leaderboard (CCKS 2023).
Huatuo-26M
Largest open Chinese medical QA resource (~26M pairs from encyclopedias, KGs, and consults); Huatuo-Lite is the usual SFT/RAG subset.
CMExam
Chinese medical licensing-exam set (~68k annotated items) used as a knowledge-recall baseline by Chinese medical and TCM LLMs.
CMB
FreedomIntelligence comprehensive Chinese medical benchmark (CMB-Exam ~280k items plus CMB-Clin cases); the most common non-TCM comparison board in TCM LLM papers.
TM-MC
KIOM literature-derived Northeast Asian medicinal-material–compound database; the 2015 release covers ~536 materials, and the 2024 2.0 paper expands to ~34k compounds. Official site currently times out.
CVDHD
Cardiovascular herbal database of 3D compound structures, targets, and pathways for virtual screening and network pharmacology. Original Peking University site currently times out.
TCMGeneDIT
Text-mined associations among TCM, genes, diseases, effects, and ingredients, with pathway and PPI links. Official NTU site no longer resolves.
TCM-Mesh
Herb–compound–gene–disease network with toxicity/side-effect records (~6,235 herbs). Official portal currently returns 403; verify via the open paper.
TCMAnalyzer
RCDD chemo-/bioinformatics service for formula/herb/ingredient networks and scaffold search (~1,493 formulas, 618 herbs). Official rcdd.org.cn currently times out.
SuperTCM
Charité biocultural TCM resource linking drugs, botanical species, ingredients, targets, KEGG pathways, and diseases (~6,516 drugs). Official tcm.charite.de no longer resolves.
CMAUP
BIDD landscape of multi-target activities, pathways, and diseases for useful plants including TCM herbs; 2024 update, downloadable from the official site.
DCABM-TCM
Literature-mined blood constituents and metabolites of TCM prescriptions and herbs, with experimental detection conditions (~1,816 structured absorbed constituents).
ITCM
Integrated formula/herb/ingredient/target platform plus 1,488 pharmacotranscriptomic profiles for 496 TCM ingredients (expression data also on Synapse).
TCMBank
Large downloadable herb–ingredient–target–disease resource with literature-mining updates after manual checks.
YaTCM
About 1,813 prescriptions, 6,220 herbs, and 47k natural products with target/pathway tools. Nankai site currently returns 403; verify via the open-access paper.
CEMTDD
Ethnic-minority herbal–compound–target–disease resource (~621 herbs, mainly Uygur/Kazakh). Original cemtdd.com now hosts something else; verify via the PMC paper.
TCMIO
Immuno-oncology TCM database of prescriptions, herbs, ingredients, targets, and pathways, with downloads and a REST API.
LTM-TCM
Symptom–prescription–plant–ingredient–target platform linking 14 source databases plus clinical and classical records (~48k formulas). Official Tasly cloud no longer resolves; verify via the paper DOI.
HIT 2.0
Manually curated herbal-ingredient–target activity pairs (~1,237 ingredients / 2,208 targets, 2000–2020 literature). Portal opens; the analysis backend port is currently down (marked site issue).
TCM Database@Taiwan
About 20k isolated-compound 2D/3D structures from 453 TCM materials for virtual screening. Original tcm.cmu.edu.tw is unreachable; verify via the PLOS paper.
BATMAN-TCM 2.0
Known and predicted TCM ingredient–target protein interactions, with greatly expanded TTI coverage and target-to-ingredient search.
SymMap 2.0
Herb–TCM symptom–modern symptom–ingredient–target–disease maps, expanded with newer pharmacopoeia records and downloadable relationship tables.
HERB 2.0
Evidence-centered TCM resource integrating clinical trials, meta-analyses, high-throughput experiments, literature, and a knowledge graph.
ETCM 2.0
Encyclopedia of TCM formulas, patent drugs, materia medica, and ingredients with target prediction and multi-scale networks; v1 site remains online.
TCM-ID
NUS BIDD formula–herb–ingredient–target resource covering pharmacopoeia, classical, and CFDA-approved prescriptions; not the same database as TCMID 2.0.
TCMID 2.0
Integrative formula–herb–ingredient–target database (distinct from NUS TCM-ID). Original megabionet site is down; verify via Zenodo extract and the NAR paper.
TCMSP
Herb–ingredient–target–disease networks with ADME parameters; public site is TCMSP 2.3 with downloadable relationship tables.
TCM-QG
About 5,000 TCM documents and 13,000 question-answer pairs from CHIP2020, for knowledge-base expansion and question generation (CC BY-SA 4.0).
TCM-PD
Yao et al. TKDE 2018 prescription topic-model set (98,334 raw / 33,765 processed symptom–herb ID pairs). CKCEST copyright, research use only. PresRecST's prescript_1195.csv is a reproduction table.
TCM-Lung
Pulmonary-disease cases from FAH-HUCM (14,948 processed; 4,484 encoded public rows of symptom/syndrome/method/prescription IDs). Full names on request. Not the same resource as TCMNSCLC.
TCM-SD
First large public TCM syndrome-differentiation text benchmark (54,152 real records, 148 syndromes, CC BY-NC-SA 4.0). Full set is in the repo folder TCM_SD_with_knowledge; Tianchi id 139034.
TCM-NER
1,997 Chinese-medicine package inserts with 59,803 entities in 13 types for building a medication knowledge graph (OpenKG / CHIP, CC BY-SA 4.0).
TCM_KG
ChatMed knowledge graph.
TCM-MKG
TCM multi-dimensional knowledge graph.
LingShu(灵枢知识图谱)
Symptom-centric contextual KG bridging TCM and biomedicine (~17.33M entities, ~39.47M relations including triples and contextual quadruples), with a portal for visualization, reasoning, and evidence-grounded QA.
OpenTCM-KG
OpenTCM gynecology classics KG (~48k entities / ~152k relations).
TCMNSCLC
Real-world NSCLC TCM reasoning dataset with fully annotated cases (pattern differentiation / treatment method / decoction / patent medicine).
ChP-TCM
KnowledgeQA and PrescriptionWriting instructions built from Chinese Pharmacopoeia Vol. I.
Traditional-Chinese-Medicine-Dataset-SFT
High-quality TCM supervised fine-tuning dataset.
TCMChat-dataset-600k
TCMChat herbal QA and recommendation instruction data (~600k).
TCM-Instruction-Tuning-ShizhenGPT
ShizhenGPT multimodal SFT data (text/vision/speech/ECG etc.; ~311k items total per paper Table 3).
ShenNong_TCM_Dataset
ShenNong TCM instruction dataset.
MedChatZH
MedChatZH TCM consultation dataset.
ChatMed_Consult_Dataset
Chinese online medical consult dataset (500k+ consults with ChatGPT replies).
CMtMedQA
ZhongJing real multi-turn doctor–patient dialogues (~70k).
Baize-TCM-Corpus-V3
~157k TCM QA items covering theory, herbs, formulas, diagnosis, acupuncture, and clinic.
neijing-sft-v1.2
~2,009 Neijing-related instruction samples for Xinghe, with thinking/output fields.
TCM-Text-Exams
Recent TCM licensure / graduate-exam text benchmark.
Medical-LLMs-Chinese-Exam
Chinese medical exam evaluation for medical LLMs.
ZhongJing-OMNI
ZhongJing-OMNI multimodal TCM eval (including tongue).
TCMEval-SDT
TCMEval-SDT: a benchmark of 300 syndrome-diagnosis cases (web, classical texts, hospital records) for evaluating TCM syndrome-differentiation reasoning, with FAIR metadata (Sci. Data 2025).
TCMBench
TCMBench: a comprehensive benchmark for evaluating LLMs in traditional Chinese medicine (arXiv 2024).
TCM-Vision-Benchmark
TCM vision benchmark (herb recognition / inspection, ~7k items).
TCM-Tongue
6,719 standardized tongue images with 20-class multi-label pathology annotations and detection baselines.
TCM-Ladder
TCM-Ladder: a multimodal QA benchmark for comprehensively evaluating TCM multimodal LLMs on real-world tasks (arXiv 2025).
TCM-Eval
Dynamic, extensible TCM evaluation platform.
TCM-BEST4SDT
Case benchmark for syndrome differentiation and treatment.
TCM-5CEval
Five-dimension deep TCM evaluation suite.
TCM-3CEval
Three-axis eval: core knowledge, classics, clinical decisions.
MTCMB
MTCMB dataset: a multi-task TCM benchmark covering knowledge, reasoning and safety, 12 subsets with ~7,100 samples (arXiv 2025).
HWTCMBench
HWTCMBench TCM capability evaluation set.
TCMEval-PA
328 multiple-choice items on prescription normative quality and safety auditing.
TCM-AQA61
Dual-view acupuncture and Tuina action-quality videos from 61 subjects each (first- and third-person), with expert categorical and continuous ratings; paired with the CME-AQA cross-view multimodal assessment framework.
LingLan
LingLan large multi-task TCM evaluation benchmark (2026).
ChiMed 2.0
Upgraded Chinese medical pretraining dataset covering TCM corpora for LLM pretraining.
classical-tcm-canon
Full-text digitizations of the TCM canon: Neijing, Nanjing, Shanghan Lun, Jingui Yaolue and warm-disease classics.
Traditional-Chinese-Medicine-Dataset-Pretrain
High-quality TCM pretraining dataset from non-Internet sources (~1GB; clinical cases, classics, encyclopedia), 99% simplified Chinese.
TCM-Pretrain-Data-ShizhenGPT
ShizhenGPT pretraining corpus (15B+ tokens reported in the paper — Stage-1 text 11.92B incl. 6.3B TCM, plus Stage-2 multimodal ~3.6B).
TCM-Ancient-Books
A corpus of nearly 700 TCM ancient-book texts.
awesome_Chinese_medical_NLP
Curated list of Chinese medical NLP resources: terminologies, corpora, word vectors, pretrained models, KGs, NER and QA (incl. CBLUE).
CPM中成药数据集
Living large-scale public Chinese patent medicine data accompanying RAG-CPMF.