xwm/SciWorld-MPO download history
xwm/SciWorld-MPO is an 8.0B-parameter reinforcement learning model by xwm. In the last 30 days it was downloaded 18 times (8 in the last 7 days), and 316 times in total.
It ranks #452,433 on the Hub by monthly downloads and #9,550 among reinforcement learning models.
It has 2 likes.
It is a fine-tune of meta-llama/Llama-3.1-8B-Instruct.
1 models build on SciWorld-MPO: 1 quantized, 0 fine-tuned, 0 adapters and 0 merges. Together with the original they were downloaded 415 times in the last 30 days. See the SciWorld-MPO galaxy.
Most downloaded derivatives
- mradermacher/SciWorld-MPO-GGUF (quantized), 397 downloads in 30 days