thu-coai/lccc download history

thu-coai/lccc is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 711 times (131 in the last 7 days), and 14,221 times in total. It ranks #25,043 among datasets by monthly downloads.

LCCC: Large-scale Cleaned Chinese Conversation corpus (LCCC) is a large corpus of Chinese conversations. A rigorous data cleaning pipeline is designed to ensure the quality of the corpus. This pipeline involves a set of rules and several classifier-based filters. Noises such as offensive or sensitiv

Open thu-coai/lccc on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.