lenML/oaast_rm_zh_jieba download history
lenML/oaast_rm_zh_jieba is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 22 times (10 in the last 7 days), and 695 times in total. It ranks #349,066 among datasets by monthly downloads.
尝试解决"llm repetition problem",使用分词模型对oaast语料进行“结巴化”数据增强,提供更强的重复内容拒绝效果。 Attempts to solve the "llm repetition problem" by using a segmentation model to enhance the oaast corpus with "stuttering" data to provide stronger rejection of duplicate content. 其次,还过滤掉了所有自我认知的微调样本。 Second, it also filters out a