themohal/saraiki-llm-dataset download history
themohal/saraiki-llm-dataset is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 212 times (91 in the last 7 days), and 337 times in total. It ranks #64,470 among datasets by monthly downloads.
Intended Uses This dataset can support: Saraiki LLM pre-training and continued pre-training Supervised fine-tuning (SFT) Instruction tuning Saraiki text generation Language understanding NLP research Low-resource language modeling Multilingual AI research LLM evaluation Academic and
Models trained on saraiki-llm-dataset
2 models list it as training data.
- themohal/saraiki-roberta-base-small-finetuned3 50 downloads in 30 days
- themohal/saraiki-qwen3-8b-cpt 37 downloads in 30 days
Open themohal/saraiki-llm-dataset on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.