HeNLP/HeDC4 download history
HeNLP/HeDC4 is a fill mask dataset on the Hugging Face Hub. In the last 30 days it was downloaded 61 times (21 in the last 7 days), and 3,384 times in total. It ranks #162,562 among datasets by monthly downloads.
Dataset Summary A Hebrew Deduplicated and Cleaned Common Crawl Corpus. A thoroughly cleaned and approximately deduplicated dataset for unsupervised learning. Citing If you use HeDC4 in your research, please cite HeRo: RoBERTa and Longformer Hebrew Language Models. @article{sha
Models trained on HeDC4
6 models list it as training data.
- HeNLP/HeRo 86 downloads in 30 days
- Slasky/HebrewGPT-1B 81 downloads in 30 days
- Slasky/HebrewGPT-1B-AdamW 27 downloads in 30 days
- Wissotsky/TamiLM-Hebrew-Nano 24 downloads in 30 days
- HeNLP/LongHeRo 23 downloads in 30 days
- omriel1/LLM2Vec-DictaLM2.0-mntp-unsup-simcse 5 downloads in 30 days
Open HeNLP/HeDC4 on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.