ClassiCC-Corpus/ClassiCC-PT download history
ClassiCC-Corpus/ClassiCC-PT is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 616 times (133 in the last 7 days), and 10,271 times in total. It ranks #28,000 among datasets by monthly downloads.
๐ ClassiCC-PT: Classified Common Crawl Corpus for Portuguese ๐ Overview ClassiCC-PT (Classified Common Crawl โ Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large languag
Models trained on ClassiCC-PT
3 models list it as training data.
- unb-labia/BERTomelo-ModernBERT-Base-v1 909 downloads in 30 days
- unb-labia/BERTomelo-ModernBERT-Large-v1 386 downloads in 30 days
- lorenzocc/NeoBERTugues 71 downloads in 30 days
Open ClassiCC-Corpus/ClassiCC-PT on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.