superdoc-dev/docx-corpus download history

superdoc-dev/docx-corpus is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 294 times (53 in the last 7 days), and 1,884 times in total. It ranks #49,911 among datasets by monthly downloads.

docx-corpus The largest classified corpus of Word documents. 736K+ .docx files from the public web, classified into 10 document types and 9 topics across 76 languages. Dataset Description This dataset contains metadata for publicly available .docx files collected from the web.

Open superdoc-dev/docx-corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.