nphearum/khmer-raw-text-3M download history
nphearum/khmer-raw-text-3M is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 18 times (2 in the last 7 days), and 498 times in total. It ranks #402,156 among datasets by monthly downloads.
Dataset Card for nphearum/khmer-raw-text-3M Dataset Summary nphearum/khmer-raw-text-3M is a large-scale raw text corpus containing approximately 50000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-tra
Models trained on khmer-raw-text-3M
3 models list it as training data.
- nphearum/Qwen3.5-4B-khmer-delta 125 downloads in 30 days
- nphearum/Gemma-4-e2b-khmer-improved 5 downloads in 30 days
- nphearum/Qwen3.5-4B-khmer-adapter – downloads in 30 days