HeNLP/HeDC4 download history

HeNLP/HeDC4 is a fill mask dataset on the Hugging Face Hub. In the last 30 days it was downloaded 61 times (21 in the last 7 days), and 3,384 times in total. It ranks #162,562 among datasets by monthly downloads.

Dataset Summary A Hebrew Deduplicated and Cleaned Common Crawl Corpus. A thoroughly cleaned and approximately deduplicated dataset for unsupervised learning. Citing If you use HeDC4 in your research, please cite HeRo: RoBERTa and Longformer Hebrew Language Models. @article{sha

Models trained on HeDC4

6 models list it as training data.

Open HeNLP/HeDC4 on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.