guychuk/benign-malicious-prompt-classification download history

guychuk/benign-malicious-prompt-classification is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 116 times (47 in the last 7 days), and 4,083 times in total. It ranks #103,237 among datasets by monthly downloads.

Important Notes This dataset goal is to help detect prompt injections / jailbreak intent. To achieve that, we decided to classify prompts to malicious only if there's an attemp to manipulate them - that means that a bad prompt (i.e asking how to create a bomb) will be classified as benign

Models trained on benign-malicious-prompt-classification

1 models list it as training data.

Open guychuk/benign-malicious-prompt-classification on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.