HebArabNlpProject/shoshan-data download history

HebArabNlpProject/shoshan-data is a token classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 74 times (13 in the last 7 days), and 290 times in total. It ranks #140,995 among datasets by monthly downloads.

Shoshan — Hebrew lemmatization data One content lemma per surface token, in context. Used to train and evaluate the Shoshan lemmatizer. file split rows train.csv / dev.csv / test.csv in-domain (Knesset + Wikipedia, IAHLT UD) 191k / 11k / 11k ood.csv (+ ood_Bagatz.csv, ood

Open HebArabNlpProject/shoshan-data on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.