jed351/Cantonese-Web-Data download history
jed351/Cantonese-Web-Data is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 56 times (10 in the last 7 days), and 884 times in total. It ranks #171,127 among datasets by monthly downloads.
Dataset Summary Cantonese has been a low-resource language in NLP. This dataset is a major step towards changing that. To our knowledge, this is the first large-scale, properly curated, and deduplicated web dataset built specifically for Cantonese. It was created by filtering years of Com
Open jed351/Cantonese-Web-Data on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.