coggpt/ParaPat download history
coggpt/ParaPat is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 9 times (4 in the last 7 days), and 496 times in total. It ranks #602,773 among datasets by monthly downloads.
This repository contains the developed parallel corpus from the open access Google Patents dataset in 74 language pairs, comprising more than 68 million sentences and 800 million tokens. Sentences were automatically aligned using the Hunalign algorithm for the largest 22 language pairs, while the ot
Open coggpt/ParaPat on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.