lordChipotle/Llama3GRPOReasoning download history
lordChipotle/Llama3GRPOReasoning is an 8.0B-parameter reinforcement learning model by lordChipotle. In the last 30 days it was downloaded 14 times (5 in the last 7 days), and 206 times in total.
It ranks #576,975 on the Hub by monthly downloads and #11,512 among reinforcement learning models.
It has 1 likes.
It is a fine-tune of meta-llama/Llama-3.1-8B-Instruct.
1 models build on Llama3GRPOReasoning: 1 quantized, 0 fine-tuned, 0 adapters and 0 merges. Together with the original they were downloaded 312 times in the last 30 days. See the Llama3GRPOReasoning galaxy.
Most downloaded derivatives
- mradermacher/Llama3GRPOReasoning-GGUF (quantized), 298 downloads in 30 days
Open lordChipotle/Llama3GRPOReasoning on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.