Locutusque/TM-DATA-V2 download history

Locutusque/TM-DATA-V2 is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 281 times (68 in the last 7 days), and 3,337 times in total. It ranks #52,003 among datasets by monthly downloads.

TM-DATA-V2 This is the dataset used to pre-train the TinyMistral-248M-V3 language model. It consists of around 10 million text documents sourced from textbooks, webpages, wikipedia, papers, and more. It is advisable to shuffle this dataset when pretraining a language model to prevent cata

Models trained on TM-DATA-V2

5 models list it as training data.

Open Locutusque/TM-DATA-V2 on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.