cis-lmu/GlotCC-V1 download history

cis-lmu/GlotCC-V1 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,900 times (227 in the last 7 days), and 119,848 times in total. It ranks #11,749 among datasets by monthly downloads.

Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com

Models trained on GlotCC-V1

3 models list it as training data.

Open cis-lmu/GlotCC-V1 on Hugging Face