Polygl0t/gigaverbo-v2-synth download history

Polygl0t/gigaverbo-v2-synth is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 546 times (84 in the last 7 days), and 3,339 times in total. It ranks #30,974 among datasets by monthly downloads.

GigaVerbo-v2 Synth: A Synthetic Dataset for Portuguese Dataset Summary GigaVerbo-v2 Synth is a large synthetic Portuguese text corpus (~9.3 billion tokens) generated to complement the GigaVerbo-v2 web-sourced dataset. Inspired by approaches like Cosmopedia, this dataset was c

Models trained on gigaverbo-v2-synth

9 models list it as training data.

Open Polygl0t/gigaverbo-v2-synth on Hugging Face