amishor/reinforce-learning-grpo download history
amishor/reinforce-learning-grpo is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 26 times (8 in the last 7 days), and 344 times in total. It ranks #306,850 among datasets by monthly downloads.
DeepSeek-R1-Reasoning-Instruct A high-quality instruction-tuning dataset derived from the official paper “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”. This dataset contains curated reasoning-focused instruction–response pairs extracted from the trai
Open amishor/reinforce-learning-grpo on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.