FredyRivera-dev/LLaDA-Sample-ES download history
FredyRivera-dev/LLaDA-Sample-ES is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 325 times (53 in the last 7 days), and 6,376 times in total. It ranks #46,117 among datasets by monthly downloads.
Dataset: LLaDA-Sample-ES Base: crscardellino/spanish_billion_words Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy maski