yuhuanstudio/c4_pretrain_zhtw download history
yuhuanstudio/c4_pretrain_zhtw is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 181 times (20 in the last 7 days), and 838 times in total. It ranks #73,018 among datasets by monthly downloads.
Dataset Card for "yuhuanstudio/c4_pretrain_zhtw" 資料集摘要 本資料集基於 C4(Colossal Clean Crawled Corpus)原始數據,並經過以下處理步驟,轉換為適用於大型語言模型(LLM)預訓練的格式: 資料清理:去除非中文內容、重複文本及不必要的 HTML 標籤,並使用pangu格式化中文語句間隔,提升語言模型的訓練品質。 格式化:將數據重新整理為適合 LLM 預訓練的結構,便於高效載入與處理。 內容說明 數據來源:Colossal Clean Crawled
Open yuhuanstudio/c4_pretrain_zhtw on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.