blue-machines/tokenwise_lid_all_scripts_dataset download history
blue-machines/tokenwise_lid_all_scripts_dataset is a token classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 48 times (6 in the last 7 days), and 106 times in total. It ranks #191,074 among datasets by monthly downloads.
blue-machines/tokenwise_lid_all_scripts_dataset Word/token-level LID training corpus (token_lid_combined_dataset.csv). Token labels are pipe-separated (|) in token_types / tokens: LID_WORD IGNORE_ENTITY Columns text token_types language language_switch tokens source
Open blue-machines/tokenwise_lid_all_scripts_dataset on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.