tvu-vlinhd11/vi-dataset-for-pretrain download history
tvu-vlinhd11/vi-dataset-for-pretrain is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 735 times (67 in the last 7 days), and 1,376 times in total. It ranks #24,449 among datasets by monthly downloads.
Dataset Card for "vi-dataset-for-pretrain" This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc. The dataset consists of: vietgpt/covid_19_news_vi hieunguyen1053/binhvq-news-corpus oscar (unshuffled_deduplicated_vi) vietgpt/wikipedia_vi
Open tvu-vlinhd11/vi-dataset-for-pretrain on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.