hathibelagal/clean_latin download history
hathibelagal/clean_latin is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 16 times (1 in the last 7 days), and 585 times in total. It ranks #433,426 among datasets by monthly downloads.
Clean Latin This is a collection of Latin strings that, hopefully, can be directly fed to an LLM during fine-tuning. I have tried to remove all unnecessary characters and strings that are non-latin. Each row in this dataset has a string with ~128 words. Citation @dataset{CleanL
Models trained on clean_latin
3 models list it as training data.
- mradermacher/llama-3.2-latin-GGUF 148 downloads in 30 days
- tensorblock/hathibelagal_llama-3.2-latin-GGUF 27 downloads in 30 days
- hathibelagal/llama-3.2-latin 24 downloads in 30 days