hathibelagal/clean_latin download history

hathibelagal/clean_latin is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 16 times (1 in the last 7 days), and 585 times in total. It ranks #433,426 among datasets by monthly downloads.

Clean Latin This is a collection of Latin strings that, hopefully, can be directly fed to an LLM during fine-tuning. I have tried to remove all unnecessary characters and strings that are non-latin. Each row in this dataset has a string with ~128 words. Citation @dataset{CleanL

Models trained on clean_latin

3 models list it as training data.

Open hathibelagal/clean_latin on Hugging Face