MultivexAI/Everyday-Language-Corpus-deduped download history
MultivexAI/Everyday-Language-Corpus-deduped is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 30 times (10 in the last 7 days), and 318 times in total. It ranks #274,414 among datasets by monthly downloads.
This dataset is a deduplicated version of the original Everyday-Language-Corpus, resulting in 7,634 unique entries. The raw dataset was processed to remove near-identical and semantically similar sentences. The deduplication was performed using embeddings generated by the BAAI/bge-small-en-v1.5 mode
Models trained on Everyday-Language-Corpus-deduped
1 models list it as training data.
- Fu01978/gpt2-mega-wiki-logic 22 downloads in 30 days
Open MultivexAI/Everyday-Language-Corpus-deduped on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.