docling-project/SynthCodeNet download history
docling-project/SynthCodeNet is an image text to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 2,908 times (1,203 in the last 7 days), and 54,730 times in total. It ranks #8,561 among datasets by monthly downloads.
SynthCodeNet SynthCodeNet is a multimodal dataset created for training the SmolDocling model. It consists of over 9.3 million synthetically generated image-text pairs, covering code snippets from 56 different programming languages. Text data was sourced from permissively licensed