UCLNLP/monoweb-dataset download history
UCLNLP/monoweb-dataset is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,253 times (50 in the last 7 days), and 10,392 times in total. It ranks #16,326 among datasets by monthly downloads.
MonoWeb Dataset MonoWeb is a multilingual pretraining corpus derived from FineWeb-Edu (English) and FineWeb2 (German, Spanish, French) by systematically removing all mixed-language documents. Released alongside the paper: The Role of Mixed-Language Documents for Multilingual Large Langua