billion-word-benchmark/lm1b download history
billion-word-benchmark/lm1b is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,432 times (196 in the last 7 days), and 147,377 times in total. It ranks #14,739 among datasets by monthly downloads.
A benchmark corpus to be used for measuring progress in statistical language modeling. This has almost one billion words in the training data.
Models trained on lm1b
2 models list it as training data.
- kuleshov-group/udlm-lm1b 1.6K downloads in 30 days
- Waspr/mdlm-lm1b-gpt2-packed – downloads in 30 days