OpenLLM-Ro/ro_sft_finepdfs download history

OpenLLM-Ro/ro_sft_finepdfs is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 299 times (155 in the last 7 days), and 3,992 times in total. It ranks #49,244 among datasets by monthly downloads.

Dataset Description FinePDFs is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Here we provide the Romanian split of FinePDFs training set, prepared for OCR: pairs of images (page

Models trained on ro_sft_finepdfs

6 models list it as training data.

Open OpenLLM-Ro/ro_sft_finepdfs on Hugging Face