lianghsun/taic-filtered download history

lianghsun/taic-filtered is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 32 times (25 in the last 7 days), and 82 times in total. It ranks #260,371 among datasets by monthly downloads.

taic-filtered:臺灣政府文件預訓練語料 本資料集來自臺灣主權 AI 訓練語料庫(TAIC),用途為繁體中文語言模型的持續預訓練。目前的 train 分割為 6,639 份文件、102,787,577 個 Gemma 4 tokens(0.102787577B)。資料蒐集截止時間沿用原版紀錄:2026-02-11 19:15:43;本版更新日期為 2026-09-30。 2026-09-29 的整理從原版 9,410 筆移出 2,178 筆字元/編碼掃描候選;2026-09-30 再從 7,232 筆中清理網址、目錄與已確認的新聞稿頁首/頁碼,並將 59

Open lianghsun/taic-filtered on Hugging Face