jsun/Prolong_64K_v2_Llama2_Tokenizer download history

jsun/Prolong_64K_v2_Llama2_Tokenizer is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 220 times (156 in the last 7 days), and 1,198 times in total. It ranks #62,930 among datasets by monthly downloads.

Prolong_64K_v2_Llama2_Tokenizer This is the Prolong_64K dataset, tokenized using the Llama-2-7b-hf tokenizer for use in Samba-style training. This dataset was used in the research paper: Rethinking Language Model Scaling under Transferable Hypersphere Optimization. The official training c

Open jsun/Prolong_64K_v2_Llama2_Tokenizer on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.