billion-word-benchmark/lm1b download history

billion-word-benchmark/lm1b is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,432 times (196 in the last 7 days), and 147,377 times in total. It ranks #14,739 among datasets by monthly downloads.

A benchmark corpus to be used for measuring progress in statistical language modeling. This has almost one billion words in the training data.

Models trained on lm1b

2 models list it as training data.

Open billion-word-benchmark/lm1b on Hugging Face