cis-lmu/GlotCC-V1 download history
cis-lmu/GlotCC-V1 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,900 times (227 in the last 7 days), and 119,848 times in total. It ranks #11,749 among datasets by monthly downloads.
Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com
Models trained on GlotCC-V1
3 models list it as training data.
- mradermacher/Alpesteibock-Llama-3-8B-Alpha-GGUF 1.1K downloads in 30 days
- kaizuberbuehler/Alpesteibock-Llama-3-8B-Alpha 21 downloads in 30 days
- aimongolia/is-mongolian 0 downloads in 30 days