FredyRivera-dev/LLaDA-Sample-10BT download history
FredyRivera-dev/LLaDA-Sample-10BT is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 872 times (82 in the last 7 days), and 7,930 times in total. It ranks #21,412 among datasets by monthly downloads.
Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Nois