vanarp/legal2023_38hrs download history

vanarp/legal2023_38hrs is an automatic speech recognition dataset on the Hugging Face Hub. In the last 30 days it was downloaded 92 times (15 in the last 7 days), and 315 times in total. It ranks #120,876 among datasets by monthly downloads.

legal2023_38hrs Court-audio ASR dataset: 38.6 h of English legal/court speech cut into per-speaker segments, with speaker-disjoint train / validation / test splits. ⚠️ Pseudo-labels, not gold. Transcripts are produced by an automatic pipeline not human annotation. Corpus WER vs an ind

Models trained on legal2023_38hrs

1 models list it as training data.

Open vanarp/legal2023_38hrs on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.