pico-lm/pretokenized-dolma download history
pico-lm/pretokenized-dolma is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,017 times (82 in the last 7 days), and 181,272 times in total. It ranks #8,310 among datasets by monthly downloads.
The Pretokenized Dolma Dataset A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library. Overview Key Features: Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a
Models trained on pretokenized-dolma
4 models list it as training data.
- pico-lm/pico-decoder-tiny 28.4K downloads in 30 days
- pico-lm/pico-decoder-small 28.3K downloads in 30 days
- pico-lm/pico-decoder-medium 27.6K downloads in 30 days
- pico-lm/pico-decoder-large 26.2K downloads in 30 days
Open pico-lm/pretokenized-dolma on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.