jbduran/bartholomew-dataset-v1 download history

jbduran/bartholomew-dataset-v1 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 948 times (156 in the last 7 days), and 5,357 times in total. It ranks #20,116 among datasets by monthly downloads.

BART Dataset v1 The first version of the BART pretraining corpus: pre-1930 English books drawn from Institutional Books 1.0 and filtered hard on OCR quality, language, date, and tokenizability. Documents 160,263 Characters 118,745,375,871 Tokens ~27B (estimated) Sha

Open jbduran/bartholomew-dataset-v1 on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.