AMR-KELEG/PTCC download history
AMR-KELEG/PTCC is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 20 times (8 in the last 7 days), and 718 times in total. It ranks #374,093 among datasets by monthly downloads.
The Parallel Tunisian Constitution Corpus (PTCC) corpus is a corpus of 149 articles written in Modern Standard Arabic and Tunisian Arabic. Tesseract was used to transform the constitution's pdf files into text files. Afterward, alignment of the parallel articles was achieved by a simple Python scrip