kashif/train_rl_agent_completions download history
kashif/train_rl_agent_completions is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 346 times (78 in the last 7 days), and 4,091 times in total. It ranks #43,945 among datasets by monthly downloads.
train_rl Completion Logs This dataset contains the on-policy generations produced during RL training with train_rl. Training details Key Value Algorithm GRPO Model (student) Qwen/Qwen3-4B-Instruct-2507 Prompt dataset VerifierEnvDataset Group size 4 Max comple
Open kashif/train_rl_agent_completions on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.