data-archetype/ocr_captions download history

data-archetype/ocr_captions is a text to image dataset on the Hugging Face Hub. In the last 30 days it was downloaded 47 times (5 in the last 7 days), and 5,446 times in total. It ranks #194,103 among datasets by monthly downloads.

OCR-Data Bucketed Captions This dataset is a bucketed WebDataset-style export of Yesianrohn/OCR-Data, with images paired with concise English captions for text-to-image training. The source dataset aggregates public OCR benchmarks with images, recognized text, text-region bounding box

Open data-archetype/ocr_captions on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.