jlli/HuCCPDF download history
jlli/HuCCPDF is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 302 times (47 in the last 7 days), and 2,747 times in total. It ranks #48,856 among datasets by monthly downloads.
HuCCPDF Approximately 113k pages of Hungarian PDFs from the Common Crawl, as featured in our paper "Synthetic Document Question Answering in Hungarian". The text field is extracted using PyMuPDF, and the ocr field is extracted using pytesseract. See other datasets from the paper: HuDocV