lhoestq/resumes-raw-pdf-for-ocr download history
lhoestq/resumes-raw-pdf-for-ocr is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 81 times (17 in the last 7 days), and 825 times in total. It ranks #132,316 among datasets by monthly downloads.
Extracted lists of pages from PDF resumes and the PDF texts. Created using this code: import io import PIL.Image from datasets import load_dataset def render(pdf): images = [] for page in pdf.pages: buffer = io.BytesIO() page.to_image(height=840).save(buffer) images.
Open lhoestq/resumes-raw-pdf-for-ocr on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.