spadeMIA/pmc_finetune_corpus_1024-2040_tokens download history
spadeMIA/pmc_finetune_corpus_1024-2040_tokens is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 70 times (51 in the last 7 days), and 174 times in total. It ranks #146,612 among datasets by monthly downloads.
PMC 1024-2040 Biomedical Fine-Tuning Corpus Summary This is a cleaned biomedical long-text corpus for autoregressive language-model fine-tuning and held-out evaluation. split rows role train 10,000 fine-tuning test 1,000 held-out evaluation The public
Models trained on pmc_finetune_corpus_1024-2040_tokens
2 models list it as training data.
- spadeMIA/pythia-1.4b-biomedical-lora-r16 14 downloads in 30 days
- spadeMIA/pythia-1.4b-biomedical-lora-r64 – downloads in 30 days
Open spadeMIA/pmc_finetune_corpus_1024-2040_tokens on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.