CodeIsAbstract/sample_sanskrit_pile download history
CodeIsAbstract/sample_sanskrit_pile is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 329 times (88 in the last 7 days), and 787 times in total. It ranks #45,656 among datasets by monthly downloads.
Sanskrit Pile v0 A tokenizer-agnostic, streamable raw pretraining corpus of Devanagari Sanskrit text, assembled for continual pretraining of Sanskrit LLMs (1B–3B proof-of-concept). Stats Documents: 1,429,519 Characters: ~6.19 Billion Approx tokens (~4 chars/tok): ~1.55