docling-project/SynthCodeNet download history

docling-project/SynthCodeNet is an image text to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 2,908 times (1,203 in the last 7 days), and 54,730 times in total. It ranks #8,561 among datasets by monthly downloads.

SynthCodeNet SynthCodeNet is a multimodal dataset created for training the SmolDocling model. It consists of over 9.3 million synthetically generated image-text pairs, covering code snippets from 56 different programming languages. Text data was sourced from permissively licensed

Open docling-project/SynthCodeNet on Hugging Face