climb-mao/Bulgarian-BabyLM download history

climb-mao/Bulgarian-BabyLM is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 18 times (5 in the last 7 days), and 347 times in total. It ranks #402,717 among datasets by monthly downloads.

Bulgarian BabyLM Dataset curated by Mila Marcheva (University of Cambridge). 28,467,275 tokens (excluding punctuation) Dataset Overview A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and

Open climb-mao/Bulgarian-BabyLM on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.