MultivexAI/Everyday-Language-Corpus-deduped download history

MultivexAI/Everyday-Language-Corpus-deduped is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 30 times (10 in the last 7 days), and 318 times in total. It ranks #274,414 among datasets by monthly downloads.

This dataset is a deduplicated version of the original Everyday-Language-Corpus, resulting in 7,634 unique entries. The raw dataset was processed to remove near-identical and semantically similar sentences. The deduplication was performed using embeddings generated by the BAAI/bge-small-en-v1.5 mode

Models trained on Everyday-Language-Corpus-deduped

1 models list it as training data.

Open MultivexAI/Everyday-Language-Corpus-deduped on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.