TrustAIRLab/HarmfulSkillBench download history
TrustAIRLab/HarmfulSkillBench is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 91 times (10 in the last 7 days), and 1,088 times in total. It ranks #122,765 among datasets by monthly downloads.
📝 Paper | 📑 arXiv | 💻 Code | 📦 Dataset HarmfulSkillBench A benchmark for evaluating LLM refusal behavior when agents are exposed to skills that describe potentially harmful capabilities. The benchmark probes whether current LLMs can detect and refuse harmful age
Open TrustAIRLab/HarmfulSkillBench on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.