BEE-spoke-data/wikipedia-20230901.en-deduped download history
BEE-spoke-data/wikipedia-20230901.en-deduped is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,566 times (175 in the last 7 days), and 35,707 times in total. It ranks #13,706 among datasets by monthly downloads.
wikipedia - 20230901.en - deduped purpose: train with less data while maintaining (most) of the quality This is really more of a "high quality diverse sample" rather than "we are trying to remove literal duplicate documents". Source dataset: graelo/wikipedia. configs
Models trained on wikipedia-20230901.en-deduped
9 models list it as training data.
- bartowski/pszemraj_jamba-900M-v0.13-KIx2-GGUF 1.3K downloads in 30 days
- BEE-spoke-data/smol_llama-101M-GQA 533 downloads in 30 days
- afrideva/smol_llama-101M-GQA-GGUF 516 downloads in 30 days
- mradermacher/Tokara-0.5B-v0.1-GGUF 248 downloads in 30 days
- BEE-spoke-data/smol_llama-81M-tied 237 downloads in 30 days
- BEE-spoke-data/mega-ar-126m-4k 226 downloads in 30 days
- pszemraj/jamba-900M-v0.13-KIx2 67 downloads in 30 days
- Kendamarron/Tokara-0.5B-v0.1 32 downloads in 30 days
- tensorblock/smol_llama-81M-tied-GGUF 19 downloads in 30 days
Open BEE-spoke-data/wikipedia-20230901.en-deduped on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.