climb-mao/Bulgarian-BabyLM download history
climb-mao/Bulgarian-BabyLM is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 18 times (5 in the last 7 days), and 347 times in total. It ranks #402,717 among datasets by monthly downloads.
Bulgarian BabyLM Dataset curated by Mila Marcheva (University of Cambridge). 28,467,275 tokens (excluding punctuation) Dataset Overview A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and
Open climb-mao/Bulgarian-BabyLM on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.