juanquivilla/sotto-transcript-cleanup download history
juanquivilla/sotto-transcript-cleanup is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 81 times (15 in the last 7 days), and 567 times in total. It ranks #133,350 among datasets by monthly downloads.
SottoASR Transcript Cleanup Dataset sotto.app · Trained Model (bf16) · MLX 5-bit Model Overview 124K+ synthetic training pairs for fine-tuning small language models on speech-to-text transcript cleanup. This dataset was used to train the SottoASR transcript cleanup m
Models trained on sotto-transcript-cleanup
1 models list it as training data.
- juanquivilla/sotto-cleanup-lfm25-350m 86 downloads in 30 days
Open juanquivilla/sotto-transcript-cleanup on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.