GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl download history

GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 39 times (10 in the last 7 days), and 836 times in total. It ranks #223,370 among datasets by monthly downloads.

Description This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to ext

Open GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.