Phonikud/phonikud-data download history
Phonikud/phonikud-data is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 644 times (120 in the last 7 days), and 3,395 times in total. It ranks #27,053 among datasets by monthly downloads.
Dataset for phonikud model The datasets contains millions of clean Hebrew sentences marked with nikud and additional phonetics marks. The format is text<TAB>phonemes Changelog (knesset_nikud - 5 millions lines) v6 Fix words with Oto such as Otobus or Otomati
Open Phonikud/phonikud-data on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.