nopperl/arxiv-image-text download history
nopperl/arxiv-image-text is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 29 times (12 in the last 7 days), and 3,653 times in total. It ranks #281,911 among datasets by monthly downloads.
arXiv Figures Dataset This dataset contains image-text pairs extracted from figures from papers published until the end of 2020 in the arXiv repository. The dataset can be used to train CLIP models. This repo contains a Parquet file containing the metadata of a WebDataset in img2dataset f
Open nopperl/arxiv-image-text on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.