mhla/pre1900-corpus download history

mhla/pre1900-corpus is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 301 times (153 in the last 7 days), and 2,657 times in total. It ranks #48,977 among datasets by monthly downloads.

Pre-1900 Corpus The training corpus for GPT-1900 — a cleaned collection of pre-1900 English-language texts with full metadata. Every document in this corpus was published before the year 1900. Schema Column Type Description text string Full document text year int64

Open mhla/pre1900-corpus on Hugging Face