lianghsun/tw-hokkien-seed-text download history
lianghsun/tw-hokkien-seed-text is a text to speech dataset on the Hugging Face Hub. In the last 30 days it was downloaded 19 times (3 in the last 7 days), and 153 times in total. It ranks #387,631 among datasets by monthly downloads.
Dataset Card for tw-hokkien-seed-text Dataset Description tw-hokkien-seed-text 是一個以台灣閩南語(台語)為主的全漢字文本資料集,專為語音合成(TTS)與自動語音辨識(ASR)模型訓練所設計。 本資料集收錄約 300 萬條台語全漢字句子,每條文字長度約 50–80 字,對應語音長度約 10–15 秒。文本內容涵蓋台灣社會各階層、各行各業的日常用語,從農漁業、宗教信仰、市井生活到現代科技,力求反映真實的台語使用情境。 所有文本均以全漢字書寫(不使用台羅拼音或漢羅混用格式),並
Open lianghsun/tw-hokkien-seed-text on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.