placeholderlabs/pretrain-encyclopedic-mix-long-context download history

placeholderlabs/pretrain-encyclopedic-mix-long-context is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 247 times (141 in the last 7 days), and 247 times in total. It ranks #57,120 among datasets by monthly downloads.

Normalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 587,625,128 (587.6M) Trainable tokens 587,625,128 (587.6M) Documents 23,631 Shards 9 UTF-8 bytes 1,978,753,989 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet

Open placeholderlabs/pretrain-encyclopedic-mix-long-context on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.