turkish-nlp-suite/BellaTurca download history
turkish-nlp-suite/BellaTurca is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,289 times (1,187 in the last 7 days), and 33,606 times in total. It ranks #7,792 among datasets by monthly downloads.
Dataset Card for BellaTurca BellaTurca is the first large-scale Turkish corpus collection for training Turkish language models. The total size is around 245GB and 30 billion words. BellaTurca's focus is high quality, diversity as well as the size. This collection is made up of five data