kipasyangin5/5m-terminal-lm-dataset download history

kipasyangin5/5m-terminal-lm-dataset is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 54 times (7 in the last 7 days), and 131 times in total. It ranks #176,012 among datasets by monthly downloads.

📦 50 Million Token Terminal & Multilingual Corpus This dataset contains the exact 51,609,269 Token raw text corpus (100.54 MB uncompressed / 32.46 MB compressed) used to train the 5M Terminal & Multilingual Language Models (kipasyangin5/5m-terminal-lm-chinchilla and kipasyangin5/5m-te

Open kipasyangin5/5m-terminal-lm-dataset on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.