mhla/pre1900-training download history

mhla/pre1900-training is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 474 times (20 in the last 7 days), and 4,455 times in total. It ranks #34,542 among datasets by monthly downloads.

Pre-1900 Training Corpus Chunked and resharded pre-1900 English text corpus, ready for language model training. Format 266 parquet shards (265 train + 1 validation) 12.8M documents (chunks of ≤8,000 characters) ~22B tokens estimated Text-only — single text column per row Row g

Open mhla/pre1900-training on Hugging Face