dataset · 2025

TCM-Pretrain-Data-ShizhenGPT

TCM-Pretrain-Data-ShizhenGPT is a dataset in the TCM LLM catalog (Awesome-TCM-LLM), dated 2025. It is maintained alongside the GitHub list of Traditional Chinese Medicine large language models.

ShizhenGPT pretraining corpus (15B+ tokens reported in the paper — Stage-1 text 11.92B incl. 6.3B TCM, plus Stage-2 multimodal ~3.6B).

Tags corpusmultimodal

Links

Related