oliverdk/adversarial-standard-Qwen2.5-32B-Instruct download history
oliverdk/adversarial-standard-Qwen2.5-32B-Instruct is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 12 times (2 in the last 7 days), and 130 times in total. It ranks #514,336 among datasets by monthly downloads.
Dataset Card for Dataset Name Adversarial gender prompts with refusal responses. Model refuses to reveal user's gender. Generated by Qwen2.5-32B-Instruct. Filtered with GPT-4.1 to remove gender leakage. Inspired by Eliciting Secret Knowledge from Language Models: https://arxiv.org/abs/2
Open oliverdk/adversarial-standard-Qwen2.5-32B-Instruct on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.