intelli-zen/language_identification download history

intelli-zen/language_identification is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 609 times (90 in the last 7 days), and 13,521 times in total. It ranks #28,249 among datasets by monthly downloads.

语种识别 Tips: 语种 zh 代表是中文, 可能是简体, 也可能是繁体. 语种 zh-cn 则代表是简体中文, zh-tw 代表繁体中文. 数据来源 数据集从网上收集整理如下: 多语言语料 数据 原始数据/项目地址 样本个数 原始数据描述 替代数据下载地址 amazon_reviews_multi Multilingual Amazon Reviews Corpus; 2010.02573 TRAIN: 1191160, VALID: 29665, TEST: 29685 我们提出了多语言亚马逊评论语料库 (MARC)

Open intelli-zen/language_identification on Hugging Face