amishor/reinforce-learning-grpo download history

amishor/reinforce-learning-grpo is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 26 times (8 in the last 7 days), and 344 times in total. It ranks #306,850 among datasets by monthly downloads.

DeepSeek-R1-Reasoning-Instruct A high-quality instruction-tuning dataset derived from the official paper “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”. This dataset contains curated reasoning-focused instruction–response pairs extracted from the trai

Open amishor/reinforce-learning-grpo on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.