PaulR11/training-runs-pairwise download history
PaulR11/training-runs-pairwise is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 323 times (295 in the last 7 days), and 3,814 times in total. It ranks #46,351 among datasets by monthly downloads.
GRPO Training Runs — Pairwise (Single-Target, With and Without FP) Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026). Four pairwise-reward runs. All use iterative pairwise comparison against a cached baseline transcript that is periodically