clankur/longcrawl64 download history

clankur/longcrawl64 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,171 times (64 in the last 7 days), and 8,684 times in total. It ranks #17,150 among datasets by monthly downloads.

LongCrawl64 is a dataset for research on architectures and algorithms for long-context modeling. It consists of 6,661,465 pre-tokenized documents, each of which is 65,536 tokens long, for a total token count of 435 billion. The dataset is preprocessed with truncation to exactly 64 KiT, shuffling alo

Open clankur/longcrawl64 on Hugging Face