anke01/ug-output download history

anke01/ug-output is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 93 times (5 in the last 7 days), and 362 times in total. It ranks #120,882 among datasets by monthly downloads.

Uyghur NLP Dataset 维吾尔语自然语言处理数据集,基于清华大学开源的 THUUyMorph 形态学标注语料库构建。 数据集 文件 条目数 说明 pretrain_corpus.jsonl 10,595 预训练语料库 morph_vocabulary.jsonl 67,832 形态学词根词典 segmentation_train.jsonl 10,572 分词训练数据 word_root_pairs.jsonl 10,391 词-词根-词缀对 格式示例 {"id"

Open anke01/ug-output on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.