blue-machines/lid_all_scripts_dataset download history

blue-machines/lid_all_scripts_dataset is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 68 times (9 in the last 7 days), and 139 times in total. It ranks #149,549 among datasets by monthly downloads.

blue-machines/lid_all_scripts_dataset All-scripts language identification corpus derived from unified_dataset.csv. Columns id text type language Stats Rows: 2,897,517 Source file: unified_dataset.csv Load from datasets import load_datas

Open blue-machines/lid_all_scripts_dataset on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.