LLM-Tuning-Safety/HEx-PHI download history
LLM-Tuning-Safety/HEx-PHI is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 340 times (75 in the last 7 days), and 9,642 times in total. It ranks #44,527 among datasets by monthly downloads.
HEx-PHI: Human-Extended Policy-Oriented Harmful Instruction Benchmark This dataset contains 330 harmful instructions (30 examples x 11 prohibited categories) for LLM harmfulness evaluation. In our work "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend
Open LLM-Tuning-Safety/HEx-PHI on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.