PleIAs/Post-OCR-Correction download history
PleIAs/Post-OCR-Correction is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,669 times (298 in the last 7 days), and 23,360 times in total. It ranks #13,098 among datasets by monthly downloads.
Post-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias. Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD0
Models trained on Post-OCR-Correction
5 models list it as training data.
- jwhong2006/t5-PostOCRAutoCorrecttion 78 downloads in 30 days
- msmth/Source-1 8 downloads in 30 days
- Stankkh/ajcrowley 0 downloads in 30 days
- ih0dl/octopus 0 downloads in 30 days
- JoftheV/Luna-Samantha – downloads in 30 days