dataset · 2025

TCM-Pretrain-Data-ShizhenGPT

ShizhenGPT pretraining corpus (15B+ tokens reported in the paper — Stage-1 text 11.92B incl. 6.3B TCM, plus Stage-2 multimodal ~3.6B).

Tags corpusmultimodal

Links

Related