SEGAgentRL/LLDS-A-GRPO-Qwen2.5-7B-Base download history
SEGAgentRL/LLDS-A-GRPO-Qwen2.5-7B-Base is a 7.6B-parameter reinforcement learning model by SEGAgentRL. In the last 30 days it was downloaded 29 times (8 in the last 7 days), and 158 times in total.
It ranks #320,759 on the Hub by monthly downloads and #4,963 among reinforcement learning models.
It has 2 likes.
It is a fine-tune of Qwen/Qwen2.5-7B.
2 models build on LLDS-A-GRPO-Qwen2.5-7B-Base: 2 quantized, 0 fine-tuned, 0 adapters and 0 merges. Together with the original they were downloaded 1,709 times in the last 30 days. See the LLDS-A-GRPO-Qwen2.5-7B-Base galaxy.
Most downloaded derivatives
- mradermacher/LLDS-A-GRPO-Qwen2.5-7B-Base-i1-GGUF (quantized), 1.2K downloads in 30 days
- mradermacher/LLDS-A-GRPO-Qwen2.5-7B-Base-GGUF (quantized), 501 downloads in 30 days