YeungNLP/firefly-pretrain-dataset download history

YeungNLP/firefly-pretrain-dataset is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 482 times (121 in the last 7 days), and 11,072 times in total. It ranks #34,104 among datasets by monthly downloads.

Firefly中文Llama2增量预训练数据 欢迎加入Firefly大模型技术交流群,关注我们的公众号。 数据简介 技术文章:QLoRA增量预训练与指令微调,及汉化Llama2的实践 该数据应为Firefly-LLaMA2-Chinese项目的增量预训练数据,一共约22GB文本,主要包含CLUE、ThucNews、CNews、COIG、维基百科等开源数据集,以及我们收集的古诗词、散文、文言文等,数据分布如下图。 模型列表 & 数据列表 我们开源了7B和13B的Base与Chat模型。Base模型是基于LLaMA2扩

Models trained on firefly-pretrain-dataset

2 models list it as training data.

Open YeungNLP/firefly-pretrain-dataset on Hugging Face