SSahas/llm_pretrain_dataset download history

SSahas/llm_pretrain_dataset is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 15 times (2 in the last 7 days), and 510 times in total. It ranks #450,867 among datasets by monthly downloads.

This is the tokenized data of salesforce/wikitext dataset. All the samples in the train set are concatenated for pretraining the llm. To see how the tokenized dataset is created please see : https://github.com/SSahas/Implementing-LLM-From-Scratch/blob/main/assets/preprocessing.ipynb PROJECT Implem

Open SSahas/llm_pretrain_dataset on Hugging Face