erhwenkuo/hh_rlhf-chinese-zhtw download history
erhwenkuo/hh_rlhf-chinese-zhtw is a reinforcement learning dataset on the Hugging Face Hub. In the last 30 days it was downloaded 64 times (15 in the last 7 days), and 1,179 times in total. It ranks #155,823 among datasets by monthly downloads.
Dataset Card for "hh_rlhf-chinese-zhtw" 此數據集合併了下列的資料: 關於有用且無害的人類偏好數據,來自 Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback。這些數據旨在為後續 RLHF 訓練訓練偏好(或獎勵)模型。這些資料不用於對話代理人的監督訓練。根據這些資料訓練對話代理可能會導致有害的模型,這種情況應該避免。 人工生成並帶註釋的紅隊對話,來自減少危害的紅隊語言模型:方法、擴展行為和經驗教訓。這些數據旨
Open erhwenkuo/hh_rlhf-chinese-zhtw on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.