HKUSTAudio/Llasa_opensource_speech_data_160k_hours_tokenized download history

HKUSTAudio/Llasa_opensource_speech_data_160k_hours_tokenized is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 795 times (155 in the last 7 days), and 9,610 times in total. It ranks #23,088 among datasets by monthly downloads.

Update (2025-02-07): Our paper has been released! This script is for merging tokenized speech datasets stored in memmap format. The input datasets can be combined to form larger training datasets. import numpy as np import os def merge_memmap_datasets(dataset_dirs, output_dir): # Ensure the

Open HKUSTAudio/Llasa_opensource_speech_data_160k_hours_tokenized on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.