pico-lm/pretokenized-dolma download history

pico-lm/pretokenized-dolma is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,017 times (82 in the last 7 days), and 181,272 times in total. It ranks #8,310 among datasets by monthly downloads.

The Pretokenized Dolma Dataset A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library. Overview Key Features: Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a

Models trained on pretokenized-dolma

4 models list it as training data.

Open pico-lm/pretokenized-dolma on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.