AnmolNimmala0/agri-slm-corpus download history
AnmolNimmala0/agri-slm-corpus is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 29 times (7 in the last 7 days), and 301 times in total. It ranks #281,348 among datasets by monthly downloads.
AgriSLM Pre-Training Corpus Domain: Agriculture (India-focused, English)Version: 1.0Total Documents: 266,691Total Tokens: ~1.32 billionFile: corpus_filtered.jsonl (6.5 GB) Purpose Pre-training corpus for a 100–300M parameter agriculture-domain language model. Covers In
Open AnmolNimmala0/agri-slm-corpus on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.