christinacdl/offensive_language_dataset download history
christinacdl/offensive_language_dataset is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 15 times (3 in the last 7 days), and 1,706 times in total. It ranks #450,867 among datasets by monthly downloads.
36.528 English texts in total, 12.955 NOT offensive and 23.573O OFFENSIVE texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label
Open christinacdl/offensive_language_dataset on Hugging Face