obswork/arxiv-ocr-benchmark-corpus download history

obswork/arxiv-ocr-benchmark-corpus is an image to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 45 times (34 in the last 7 days), and 1,042 times in total. It ranks #201,034 among datasets by monthly downloads.

arxiv-ai-ml-images Rasterized page images for the arXiv AI/ML OCR benchmark corpus. Pages are rendered at 144 DPI, encoded as WebP (quality=85, method=6), and packed into parquet shards with the Hugging Face Image feature so datasets.load_dataset decodes them automatically. Source PDFs: t

Open obswork/arxiv-ocr-benchmark-corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.