RamAnanth1/toolrl-rlla4k download history
RamAnanth1/toolrl-rlla4k is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 35 times (8 in the last 7 days), and 967 times in total. It ranks #242,743 among datasets by monthly downloads.
ToolRL rlla_4k A 4,000-example dataset for training tool-using LLM agents with reinforcement learning. This is the processed RL training split released by the ToolRL project for the paper: ToolRL: Reward is All Tool Learning Needs The dataset is intended for: GRPO PPO RLHF / RLVR tool /
Models trained on toolrl-rlla4k
1 models list it as training data.
- RamAnanth1/qwen2.5-1.5b-grpo-tool-calling 21 downloads in 30 days