MongoDB/subset_arxiv_papers_with_embeddings download history
MongoDB/subset_arxiv_papers_with_embeddings is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 2,462 times (122 in the last 7 days), and 34,354 times in total. It ranks #9,666 among datasets by monthly downloads.
This dataset is a curated subset of the original arXiv dataset, each entry enriched with a 256-dimensional embedding vector. The embeddings are generated using OpenAI's "text-embedding-3-small" model. For each data point, the embedding is created by concatenating the text of the title, author(s), an
Open MongoDB/subset_arxiv_papers_with_embeddings on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.