placeholderlabs/pretrain-web-mix-long-context download history
placeholderlabs/pretrain-web-mix-long-context is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 541 times (60 in the last 7 days), and 541 times in total. It ranks #31,202 among datasets by monthly downloads.
Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 8,689,580,607 (8.7B) Trainable tokens 8,689,580,607 (8.7B) Documents 281,846 Shards 89 UTF-8 bytes 37,540,769,483 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parqu
Open placeholderlabs/pretrain-web-mix-long-context on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.