AlienKevin/hkcancor-multi download history

AlienKevin/hkcancor-multi is a token classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 37 times (3 in the last 7 days), and 889 times in total. It ranks #232,531 among datasets by monthly downloads.

This data is the subset of the Hong Kong Cantonese Corpus (HKCanCor) that has been re-segmented by the multi-tiered word segmentation scheme described in the following paper: Charles Lam, Chaak-ming Lau, and Jackson L. Lee. 2024. Multi-Tiered Cantonese Word Segmentation. In Proceedings of the 2024 J

Models trained on hkcancor-multi

4 models list it as training data.

Open AlienKevin/hkcancor-multi on Hugging Face