devngho/the_stack_llm_annotations download history
devngho/the_stack_llm_annotations is a text classification dataset on the Hugging Face Hub. In the last 30 days it was downloaded 16 times (3 in the last 7 days), and 447 times in total. It ranks #433,426 among datasets by monthly downloads.
Dataset 이 데이터셋은 fineweb-edu의 방법을 여러 프로그래밍 언어에 적용하기 위해 만들어진 합성 데이터셋입니다. 기존에 존재하던 HuggingFaceTB/smollm-corpus의 Python-edu는 Python으로만 한정되어 있었습니다. 이 데이터셋은 bigcode/the-stack-dedup에서 21개의 프로그래밍 언어에서 각각 30k 샘플을 추출해 평가해 여러 언어에 대응합니다. 구체적으로는 devngho/the-stack-mini-nonshuffled의 첫 30k 샘플이 사용되었습니다. T
Models trained on the_stack_llm_annotations
1 models list it as training data.
- devngho/code_edu_classifier_v2_microsoft_codebert-base 0 downloads in 30 days
Open devngho/the_stack_llm_annotations on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.