yhavinga/ccmatrix download history

yhavinga/ccmatrix is a translation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 425 times (47 in the last 7 days), and 784,865 times in total. It ranks #37,595 among datasets by monthly downloads.

CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WEB We show that margin-based bitext mining in LASER's multilingual sentence space can be applied to monolingual corpora of billions of sentences to produce high quality aligned translation data. We use thirty-two snapshots of a cu

Models trained on ccmatrix

8 models list it as training data.

Open yhavinga/ccmatrix on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.