siddharthmb/2026.RA.Fairness-GRPO-v2 download history

siddharthmb/2026.RA.Fairness-GRPO-v2 is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 34 times (17 in the last 7 days), and 205 times in total. It ranks #248,325 among datasets by monthly downloads.

2026.RA.Fairness-GRPO-v2 — the complete λ-frontier of a fairness-trained LLM negotiator What this is. The full evaluation record of the fairness-GRPO v2 campaign (experiments/rational_agents/ in the ii_mats repo): GRPO training of Qwen3-8B (LoRA r32/α64) on an engine-computed, text-bl

Models trained on 2026.RA.Fairness-GRPO-v2

1 models list it as training data.

Open siddharthmb/2026.RA.Fairness-GRPO-v2 on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.