GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl download history
GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 39 times (10 in the last 7 days), and 836 times in total. It ranks #223,370 among datasets by monthly downloads.
Description This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to ext
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.