XiaSheng/FreeChunk-corpus download history

XiaSheng/FreeChunk-corpus is a text retrieval dataset on the Hugging Face Hub. In the last 30 days it was downloaded 82 times (24 in the last 7 days), and 287 times in total. It ranks #131,143 among datasets by monthly downloads.

FreeChunk Corpus This dataset is derived from The Pile and is used for evaluating and training chunking models, specifically for the FreeChunk framework. Dataset Structure The dataset consists of documents split into sentences. Features uuid: Unique identifier for t

Open XiaSheng/FreeChunk-corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.