nopperl/arxiv-image-text download history

nopperl/arxiv-image-text is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 29 times (12 in the last 7 days), and 3,653 times in total. It ranks #281,911 among datasets by monthly downloads.

arXiv Figures Dataset This dataset contains image-text pairs extracted from figures from papers published until the end of 2020 in the arXiv repository. The dataset can be used to train CLIP models. This repo contains a Parquet file containing the metadata of a WebDataset in img2dataset f

Open nopperl/arxiv-image-text on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.