TCM LLM · Datasets

TCM LLM datasets

Public datasets for training and evaluating TCM LLMs: instruction data, case QA, formulas, classical corpora, and benchmarks.

Maintained by Yang Tan · synced with Awesome-TCM-LLM · 39 entries

dataset · 2026.07

TCM-RobustSDT

TCM-RobustSDT: a robustness benchmark dataset for LLM clinical reasoning in TCM (Figshare).

dataset · 2026.05

HSQ-TD(健身气功指令微调数据集)

HSQ-TD: the first instruction-tuning dataset for health Qigong/wellness, with 57,843 instructions distilled from official textbooks and professional literature (ScienceDB).

dataset · 2025

TCM_KG

ChatMed knowledge graph.

dataset · 2025

TCM-MKG

TCM multi-dimensional knowledge graph.

dataset · 2025

OpenTCM-KG

OpenTCM gynecology classics KG (~48k entities / ~152k relations).

dataset · 2026 · Zenodo

TCMNSCLC

Real-world NSCLC TCM reasoning dataset with fully annotated cases (pattern differentiation / treatment method / decoction / patent medicine).

dataset · 2024

ChP-TCM

KnowledgeQA and PrescriptionWriting instructions built from Chinese Pharmacopoeia Vol. I.

dataset · 2025

Traditional-Chinese-Medicine-Dataset-SFT

High-quality TCM supervised fine-tuning dataset.

dataset · 2025

TCMChat-dataset-600k

TCMChat herbal QA and recommendation instruction data (~600k).

dataset · 2025

TCM-Instruction-Tuning-ShizhenGPT

ShizhenGPT multimodal SFT data (text/vision/speech/ECG etc.; ~311k items total per paper Table 3).

dataset · 2025

ShenNong_TCM_Dataset

ShenNong TCM instruction dataset.

dataset · 2025

MedChatZH

MedChatZH TCM consultation dataset.

dataset · 2025

ChatMed_Consult_Dataset

Chinese online medical consult dataset (500k+ consults with ChatGPT replies).

dataset · 2025

CMtMedQA

ZhongJing real multi-turn doctor–patient dialogues (~70k).

dataset · 2025

Baize-TCM-Corpus-V3

~157k TCM QA items covering theory, herbs, formulas, diagnosis, acupuncture, and clinic.

dataset · 2026

neijing-sft-v1.2

~2,009 Neijing-related instruction samples for Xinghe, with thinking/output fields.

dataset · 2025

TCM-Text-Exams

Recent TCM licensure / graduate-exam text benchmark.

dataset · 2025

Medical-LLMs-Chinese-Exam

Chinese medical exam evaluation for medical LLMs.

dataset · 2025

ZhongJing-OMNI

ZhongJing-OMNI multimodal TCM eval (including tongue).

dataset · 2025 · Scientific Data

TCMEval-SDT

TCMEval-SDT: a benchmark of 300 syndrome-diagnosis cases (web, classical texts, hospital records) for evaluating TCM syndrome-differentiation reasoning, with FAIR metadata (Sci. Data 2025).

dataset · 2025

TCMBench

TCMBench: a comprehensive benchmark for evaluating LLMs in traditional Chinese medicine (arXiv 2024).

dataset · 2025

TCM-Vision-Benchmark

TCM vision benchmark (herb recognition / inspection, ~7k items).

dataset · 2025

TCM-Tongue

6,719 standardized tongue images with 20-class multi-label pathology annotations and detection baselines.

dataset · 2025 · arXiv

TCM-Ladder

TCM-Ladder: a multimodal QA benchmark for comprehensively evaluating TCM multimodal LLMs on real-world tasks (arXiv 2025).

dataset · 2025

TCM-Eval

Dynamic, extensible TCM evaluation platform.

dataset · 2025

TCM-BEST4SDT

Case benchmark for syndrome differentiation and treatment.

dataset · 2025

TCM-5CEval

Five-dimension deep TCM evaluation suite.

dataset · 2025

TCM-3CEval

Three-axis eval: core knowledge, classics, clinical decisions.

dataset · 2025

MTCMB

MTCMB dataset: a multi-task TCM benchmark covering knowledge, reasoning and safety, 12 subsets with ~7,100 samples (arXiv 2025).

dataset · 2025

HWTCMBench

HWTCMBench TCM capability evaluation set.

dataset · 2026

TCMEval-PA

328 multiple-choice items on prescription normative quality and safety auditing.

dataset · 2026

LingLan

LingLan large multi-task TCM evaluation benchmark (2026).

dataset · 2025

ChiMed 2.0

Upgraded Chinese medical pretraining dataset covering TCM corpora for LLM pretraining.

dataset · 2025

classical-tcm-canon

Full-text digitizations of the TCM canon: Neijing, Nanjing, Shanghan Lun, Jingui Yaolue and warm-disease classics.

dataset · 2025

Traditional-Chinese-Medicine-Dataset-Pretrain

High-quality TCM pretraining dataset from non-Internet sources (~1GB; clinical cases, classics, encyclopedia), 99% simplified Chinese.

dataset · 2025

TCM-Pretrain-Data-ShizhenGPT

ShizhenGPT pretraining corpus (15B+ tokens reported in the paper — Stage-1 text 11.92B incl. 6.3B TCM, plus Stage-2 multimodal ~3.6B).

dataset · 2025

TCM-Ancient-Books

A corpus of nearly 700 TCM ancient-book texts.

dataset · 2025

awesome_Chinese_medical_NLP

Curated list of Chinese medical NLP resources: terminologies, corpora, word vectors, pretrained models, KGs, NER and QA (incl. CBLUE).

dataset · 2025

CPM中成药数据集

Living large-scale public Chinese patent medicine data accompanying RAG-CPMF.