wenge-research/yayi2_pretrain_data download history
wenge-research/yayi2_pretrain_data is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 4,231 times (161 in the last 7 days), and 46,120 times in total. It ranks #6,488 among datasets by monthly downloads.
介绍/Introduction 本数据集源自雅意训练语料,我们精选了约100B数据,数据大小约为500GB。我们期望通过雅意预训练数据的开源推动中文预训练大模型开源社区的发展,并积极为此贡献力量。通过开源,我们与每一位合作伙伴共同构建雅意大模型生态。 We opensource the pre-trained dataset in this release, it should contain more than 100B tokens depending on the tokenizer you use, requiring more than 500GB of loc
Spaces using yayi2_pretrain_data
- Model Pulse 8 likes