FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT download history

FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,422 times (177 in the last 7 days), and 10,715 times in total. It ranks #14,819 among datasets by monthly downloads.

📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale

Models trained on TCM-Pretrain-Data-ShizhenGPT

14 models list it as training data.

Open FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.