The xinlai/DeepSeekMath-RL-Step-DPO galaxy
1 models are built on xinlai/DeepSeekMath-RL-Step-DPO: 1 quantized, 0 fine-tuned, 0 adapters and 0 merges. Together they were downloaded 1,431 times in the last 30 days.
Most downloaded direct derivatives
- mradermacher/DeepSeekMath-RL-Step-DPO-GGUF (quantized), 1.4K downloads in 30 days