stindardlogic/helpfulness-safety-calibration-dpo-100k download history

stindardlogic/helpfulness-safety-calibration-dpo-100k is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 34 times (6 in the last 7 days), and 122 times in total. It ranks #248,758 among datasets by monthly downloads.

Helpfulness-Safety Calibration DPO (100K) 100,000 DPO preference pairs for calibrating the helpfulness-safety tradeoff in language models. Each example contains a prompt, a chosen response (correct handling), and a rejected response (incorrect handling) — covering both over-refusal an

Open stindardlogic/helpfulness-safety-calibration-dpo-100k on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.