zeyuzy/qwen3-0.6b-beavertails-reward-norm download history

zeyuzy/qwen3-0.6b-beavertails-reward-norm is a reinforcement learning dataset on the Hugging Face Hub. In the last 30 days it was downloaded 40 times (18 in the last 7 days), and 120 times in total. It ranks #219,113 among datasets by monthly downloads.

Qwen3-0.6B BeaverTails Reward — Normalized (sigmoid / min-max) 两通道 reward 数组:通道 0 = safety,通道 1 = utility。 safety 通道本身已在 [0,1],各版本均不改动;utility 通道分别用 sigmoid 或全局 min-max 归一化。 形状:(6000, 90, 2),dtype float32 sigmoid:1 / (1 + exp(-u)) min-max:(u - u.min()) / (u.max() - u.min())

Open zeyuzy/qwen3-0.6b-beavertails-reward-norm on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.