ahmed-farhanur-rashid/bn-foundational-pretrain-corpus download history

ahmed-farhanur-rashid/bn-foundational-pretrain-corpus is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 943 times (100 in the last 7 days), and 2,130 times in total. It ranks #20,229 among datasets by monthly downloads.

Bangla Pretraining Dataset This dataset is a large-scale deduplicated and unicode normalized corpus designed for pretraining Bangla language models. It combines datasets from TituLLM Bangla CC, FineWeb-Edu, and parallel translation corpora from NLLB and BanglaNMT. Order

Models trained on bn-foundational-pretrain-corpus

5 models list it as training data.

Open ahmed-farhanur-rashid/bn-foundational-pretrain-corpus on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.