CatholicCorpus/catholiccorpus-text download history
CatholicCorpus/catholiccorpus-text is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 227 times (84 in the last 7 days), and 22,201 times in total. It ranks #60,993 among datasets by monthly downloads.
CatholicCorpus — Extracted Text Pre-extracted plain text from the CatholicCorpus — 2,000 years of the Catholic intellectual tradition, ready for NLP, RAG, and digital humanities. This dataset contains 47,407 plain text files (5.7 GB, 2.64 billion GPT-2 tokens) extracted from the raw sourc
Open CatholicCorpus/catholiccorpus-text on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.