japanese-data-analyze/japanese-tiny-llm-10m-artifacts download history

japanese-data-analyze/japanese-tiny-llm-10m-artifacts is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 517 times (517 in the last 7 days), and 517 times in total. It ranks #32,289 among datasets by monthly downloads.

JMicro-10M tokenizer and pilot data handoff 日本語約10M言語モデルのPhase 4学習へ移行するための固定成果物です。 4k/8k SentencePiece Unigramトークナイザーと、前処理・split済みpilotデータを共有します。 モデル重みは含みません。本学習はまだ実施していません。 Path Contents artifacts/phase2/4096/, artifacts/phase2/8192/ Tokenizer model/vocabulary, wrapper/spe

Open japanese-data-analyze/japanese-tiny-llm-10m-artifacts on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.