PuristanLabs1/urdu-ocr-1M download history

PuristanLabs1/urdu-ocr-1M is an image to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 177 times (16 in the last 7 days), and 1,594 times in total. It ranks #74,238 among datasets by monthly downloads.

Urdu OCR Dataset (1.5 Million Samples) Dataset Summary This is a large-scale synthetic dataset for Urdu Optical Character Recognition (OCR), featuring a groundbreaking Nastaliq collection and a robust Naskh base. Nastaliq (Primary): 499,845 samples rendered with authentic Jame

Models trained on urdu-ocr-1M

1 models list it as training data.

Open PuristanLabs1/urdu-ocr-1M on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.