datalama/pretrain-nllb-filtered download history

datalama/pretrain-nllb-filtered is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 3,805 times (181 in the last 7 days), and 20,010 times in total. It ranks #6,985 among datasets by monthly downloads.

pretrain-nllb-filtered Filtered parallel corpus from allenai/nllb for cross-lingual embedding pretraining. Schema {"query": "string", "pos": ["string", ...]} query: source language sentence pos: target language sentence(s) Configs (51 language pairs) Con

Open datalama/pretrain-nllb-filtered on Hugging Face