erhwenkuo/rlhf_reward_single_round-chinese-zhtw download history

erhwenkuo/rlhf_reward_single_round-chinese-zhtw is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 20 times (5 in the last 7 days), and 1,284 times in total. It ranks #374,650 among datasets by monthly downloads.

Dataset Card for "rlhf_reward_single_round-chinese-zhtw" 基於 anthropic 的 Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback 論文開源的關於有助和無害的人類偏好資料。 這些數據旨在為後續的 RLHF 訓練訓練偏好(或獎勵)模型。 來源資料集 本資料集來自 beyond/rlhf-reward-single-round-trans_chinese, 并使用

Open erhwenkuo/rlhf_reward_single_round-chinese-zhtw on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.