declare-lab/CategoricalHarmfulQA download history

declare-lab/CategoricalHarmfulQA is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 746 times (284 in the last 7 days), and 9,799 times in total. It ranks #24,188 among datasets by monthly downloads.

CatQA: A categorical harmful questions dataset CatQA is used in LLM safety realignment research: Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic (Paper, Code) How to download from datasets import load_dataset d

Open declare-lab/CategoricalHarmfulQA on Hugging Face