magyar-nlp-szine-java/europarl_hun download history

magyar-nlp-szine-java/europarl_hun is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 10 times (1 in the last 7 days), and 119 times in total. It ranks #569,006 among datasets by monthly downloads.

The Europarl Hungarian corpus extracted from the proceedings of the European Parliament. (https://www.statmt.org/europarl/) Special HTML entities are removed from the data. We split the raw text files into segments of 2,048 tokens. tokens=17,789,186 words: 12,606,986 sentences: 658,824

Models trained on europarl_hun

1 models list it as training data.

Open magyar-nlp-szine-java/europarl_hun on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.