bigcode/the-stack-smol download history

bigcode/the-stack-smol is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 25,334 times (7,566 in the last 7 days), and 247,561 times in total. It ranks #1,219 among datasets by monthly downloads.

Dataset Description A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the original dataset. The dataset has 2.6GB of text (code). Languages The dataset contains 30 programming languages: "assembly", "batchfile", "c++", "c

Models trained on the-stack-smol

22 models list it as training data.

Spaces using the-stack-smol

Open bigcode/the-stack-smol on Hugging Face