guychuk/benign-malicious-prompt-classification download history
guychuk/benign-malicious-prompt-classification is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 116 times (47 in the last 7 days), and 4,083 times in total. It ranks #103,237 among datasets by monthly downloads.
Important Notes This dataset goal is to help detect prompt injections / jailbreak intent. To achieve that, we decided to classify prompts to malicious only if there's an attemp to manipulate them - that means that a bad prompt (i.e asking how to create a bomb) will be classified as benign
Models trained on benign-malicious-prompt-classification
1 models list it as training data.
- AminaAkhtar/Llama-3.2-1B-prompt-classifier 19 downloads in 30 days
Open guychuk/benign-malicious-prompt-classification on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.