zhaohq/PureRL-1.5B-v7-stage1-reasoning download history
zhaohq/PureRL-1.5B-v7-stage1-reasoning is a 1.8B-parameter text generation model by zhaohq. In the last 30 days it was downloaded 20 times (3 in the last 7 days), and 307 times in total.
It ranks #411,719 on the Hub by monthly downloads and #123,174 among text generation models.
It has 0 likes.
It is a fine-tune of Qwen/Qwen2.5-Math-1.5B.
7 models build on PureRL-1.5B-v7-stage1-reasoning: 0 quantized, 7 fine-tuned, 0 adapters and 0 merges. Together with the original they were downloaded 117 times in the last 30 days. See the PureRL-1.5B-v7-stage1-reasoning galaxy.
Most downloaded derivatives
- zhaohq/PureRL-1.5B-v7-s2-l2-kl-w3-b0 (finetune), 16 downloads in 30 days
- zhaohq/PureRL-1.5B-v7-s2-l2-kl-w2-b0 (finetune), 15 downloads in 30 days
- zhaohq/PureRL-1.5B-v7-s2-l2-kl-w3-b1 (finetune), 15 downloads in 30 days
- zhaohq/PureRL-1.5B-v7-s2-l2-kl-w0-b0 (finetune), 14 downloads in 30 days
- zhaohq/PureRL-1.5B-v7-s2-l2-kl-w2-b1 (finetune), 13 downloads in 30 days
- zhaohq/PureRL-1.5B-v7-s2-l2-kl-w0-b1 (finetune), 12 downloads in 30 days
- zhaohq/PureRL-1.5B-v7-s2-l2-kl-w1-b1 (finetune), 12 downloads in 30 days