pixparse/pdfa-eng-wds download history

pixparse/pdfa-eng-wds is an image to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 7,701 times (1,095 in the last 7 days), and 156,657 times in total. It ranks #3,865 among datasets by monthly downloads.

Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset

Models trained on pdfa-eng-wds

17 models list it as training data.

Spaces using pdfa-eng-wds

Open pixparse/pdfa-eng-wds on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.