kashif/train_rl_agent_completions download history

kashif/train_rl_agent_completions is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 346 times (78 in the last 7 days), and 4,091 times in total. It ranks #43,945 among datasets by monthly downloads.

train_rl Completion Logs This dataset contains the on-policy generations produced during RL training with train_rl. Training details Key Value Algorithm GRPO Model (student) Qwen/Qwen3-4B-Instruct-2507 Prompt dataset VerifierEnvDataset Group size 4 Max comple

Open kashif/train_rl_agent_completions on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.