PleIAs/Post-OCR-Correction download history

PleIAs/Post-OCR-Correction is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,669 times (298 in the last 7 days), and 23,360 times in total. It ranks #13,098 among datasets by monthly downloads.

Post-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias. Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD0

Models trained on Post-OCR-Correction

5 models list it as training data.

Open PleIAs/Post-OCR-Correction on Hugging Face