mdnaseif/hafith-synthetic-1m download history

mdnaseif/hafith-synthetic-1m is an image to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 410 times (103 in the last 7 days), and 2,235 times in total. It ranks #38,654 among datasets by monthly downloads.

HAFITH Synthetic Dataset (1M Samples) Synthetic dataset of 1 million manuscript-style Arabic text line images for training the HAFITH OCR model. Dataset Summary Total Samples: 1,000,000 (900K train / 50K val / 50K test) Text Source: ArabicText-Large (244M words) Fonts: 350 Ara

Models trained on hafith-synthetic-1m

1 models list it as training data.

Open mdnaseif/hafith-synthetic-1m on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.