data-archetype/ocr_captions download history
data-archetype/ocr_captions is a text to image dataset on the Hugging Face Hub. In the last 30 days it was downloaded 47 times (5 in the last 7 days), and 5,446 times in total. It ranks #194,103 among datasets by monthly downloads.
OCR-Data Bucketed Captions This dataset is a bucketed WebDataset-style export of Yesianrohn/OCR-Data, with images paired with concise English captions for text-to-image training. The source dataset aggregates public OCR benchmarks with images, recognized text, text-region bounding box
Open data-archetype/ocr_captions on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.