HAERAE-HUB/KOREAN-SyntheticText-1.5B download history

HAERAE-HUB/KOREAN-SyntheticText-1.5B is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 253 times (27 in the last 7 days), and 4,743 times in total. It ranks #56,059 among datasets by monthly downloads.

KOREAN-SyntheticText KOREAN-SyntheticText is a successor of the KOREAN-WEBTEXT project in our mission to create high-quality Korean corpora. The dataset consists of 1.4B tokens generated over 600 H100 hours following the Cosmopedia project. The dataset has been generated using a 100B + op

Models trained on KOREAN-SyntheticText-1.5B

2 models list it as training data.

Open HAERAE-HUB/KOREAN-SyntheticText-1.5B on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.