tvu-vlinhd11/vi-dataset-for-pretrain download history

tvu-vlinhd11/vi-dataset-for-pretrain is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 735 times (67 in the last 7 days), and 1,376 times in total. It ranks #24,449 among datasets by monthly downloads.

Dataset Card for "vi-dataset-for-pretrain" This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc. The dataset consists of: vietgpt/covid_19_news_vi hieunguyen1053/binhvq-news-corpus oscar (unshuffled_deduplicated_vi) vietgpt/wikipedia_vi

Open tvu-vlinhd11/vi-dataset-for-pretrain on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.