MartinThoma/wili_2018 download history
MartinThoma/wili_2018 is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 650 times (101 in the last 7 days), and 29,486 times in total. It ranks #26,877 among datasets by monthly downloads.
Dataset Card for wili_2018 Dataset Summary WiLI-2018, the Wikipedia language identification benchmark dataset, contains 235000 paragraphs of 235 languages. The dataset is balanced and a train-test split is provided. Supported Tasks and Leaderboards [Needs More Inform