cloudaocr/clouda-ocr-canonical-v1 download history
cloudaocr/clouda-ocr-canonical-v1 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 5,184 times (5,184 in the last 7 days), and 5,184 times in total. It ranks #5,543 among datasets by monthly downloads.
Clouda OCR canonical v1 The cleanup-approved training corpus contains 28,691 procedural synthetic pages with stable IDs, unchanged images and GT, source provenance, and hashes. Use manifests/train_approved.parquet and the three canonical_v1/shards/train_approved/train-*.tar shards. Ea