nphearum/khmer-raw-text-3M-v2 download history
nphearum/khmer-raw-text-3M-v2 is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 69 times (20 in the last 7 days), and 372 times in total. It ranks #148,091 among datasets by monthly downloads.
Dataset Card for nphearum/khmer-raw-text-3M-v2 Dataset Summary nphearum/khmer-raw-text-3M-v2 is a large-scale raw text corpus containing approximately 200_000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued