rubentium/sparseatt-fineweb-large-tokenized download history

rubentium/sparseatt-fineweb-large-tokenized is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 78 times (78 in the last 7 days), and 78 times in total. It ranks #135,785 among datasets by monthly downloads.

SparseAtt FineWeb Large Tokenized Export This repository contains the pretokenized training export used by SparseAtt. The source corpus is FineWeb, distributed under ODC-By 1.0; downstream users must also follow the applicable Common Crawl terms. The full export has 12,748,481 records

Open rubentium/sparseatt-fineweb-large-tokenized on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.