RLAIF on Hugging Face: downloads and rankings
RLAIF has 12 tracked models, downloaded 29 times in the last 30 days and 202 times in total; 135 datasets, downloaded 3,237 times in the last 30 days. RLAIF Wrapped: the last 12 months.
Models by downloads in the last 30 days
- RLAIF/reward-model-grpo 8 downloads, model
- RLAIF/llama-3b-open-r1-50k-sft 8 downloads, model
- RLAIF/Qwen3-1.7B_grpo_lr2e-7_n4_step30 4 downloads, model
- RLAIF/15-w-error-masking-temp-0-verifier-in-context-train-in-context-inference-8-model 3 downloads, model
- RLAIF/grpo_thinking_ultrafeedback-original_32_64_4_3e-3_2e-7_step-120_1.7B 3 downloads, model
- RLAIF/grpo_5e-7_4_1.7B-best 3 downloads, model
- RLAIF/22-sequential-temp-0-verifier-no-best-oracle-in-context-train-8 0 downloads, model
- RLAIF/sft-external 0 downloads, text generation
- RLAIF/sft-llama-3.1-8b-external 0 downloads, text generation
- RLAIF/22-sequential-temp-0-verifier-oracle-in-context-train-8-w-error-masking 0 downloads, model
- RLAIF/sft-gemma-2-9b-base-sft-llama-405b-instruct-correct-only-format-lr-5e-06-bs-64 0 downloads, text generation
- RLAIF/sft-llama8b-prm-800k-correct-only 0 downloads, text generation
Datasets by downloads in the last 30 days
- RLAIF/numina-math-llama-3.1-8b-bon-meta-cot 622 downloads
- RLAIF/optim_policy_pretrain-pythia-160m_lr0.0001_bs24_wp1_wd0.01_ep0_cp35k-merged 416 downloads
- RLAIF/pretext-ui-harbor-runs-v0 355 downloads
- RLAIF/WritingPrompts-Filtered 107 downloads
- RLAIF/mbpp 73 downloads
- RLAIF/dpo_answer_openorca_base_nathan_2e-6_0.02_1.7B_4B_with_gold_labels_kl_estimation 71 downloads
- RLAIF/dpo_thinking_with_gold_labels_kl_estimation 62 downloads
- RLAIF/iGSM-1M-retry0.0 52 downloads
- RLAIF/dpo_answer_openorca_skywork_rejudged_filtered_1e-6_0.02_1.7B_4B_with_gold_labels_kl_estimation 52 downloads
- RLAIF/iGSM-1M-retry0.5 49 downloads
- RLAIF/dpo_answer_offtheshelf_openorca_1e-6_0.02_0.6B_0.6B_with_gold_labels_kl_estimation 48 downloads
- RLAIF/TIR-Batched-PRM-Seed-Rollouts 31 downloads
- RLAIF/genrm-uf-qwen3-4b-angel-judge-qwen-3-32b-thinking-jt07-n113978 30 downloads
- RLAIF/dpo_thinking_ultrafeedback_rejudged_openorca_0.02_with_gold_labels_kl_estimation 28 downloads
- RLAIF/Value-v1-NUMINA-V1-Blocks-Merged-2964-problems-step-len-filtered 27 downloads
- RLAIF/iGSM-1M-retry0.6 27 downloads
- RLAIF/dpo_answer_ultrafeedback_rejudged_openorca_0.02_with_gold_labels_kl_estimation 26 downloads
- RLAIF/webgpt 26 downloads
- RLAIF/dpo_uf_rejudged_mixed_openorca_with_gold_labels_kl_estimation 23 downloads
- RLAIF/STAR-TRAIN-math_lama-star-iter4 22 downloads
- RLAIF/dpo_answer_base_openorca_0.02_with_gold_labels_kl_estimation 22 downloads
- RLAIF/dpo_answer_ultrainteract_openorca_0.02_with_gold_labels_kl_estimation 21 downloads
- RLAIF/dpo_uf_rejudged_mixed_openorca_kl_estimation 19 downloads
- RLAIF/dpo_answer_only_0.05_with_gold_labels_kl_estimation 19 downloads
- RLAIF/dpo_answer_openorca_openorca_argilla_improved_1e-6_0.02_1.7B_4B_with_gold_labels_kl_estimation 18 downloads
- RLAIF/math 18 downloads
- RLAIF/Value-v1-NUMINA-V1-Blocks-Merged 18 downloads
- RLAIF/dpo_answer_openorca_helpsteer3_improved_1e-6_0.02_1.7B_4B_with_gold_labels_kl_estimation 17 downloads
- RLAIF/dpo_thinking_base_openorca_0.02_1.7B-4B_with_gold_labels_kl_estimation 17 downloads
- RLAIF/dpo_answer_openorca_baseline_mix_1e-6_0.02_1.7B_4B_with_gold_labels_kl_estimation 17 downloads
- RLAIF/Value-v2-NUMINA-V2-Blocks-Merged-1999-problems-step-len-filtered 17 downloads
- RLAIF/dpo_answer_openorca_angel_nathan_1e-6_0.02_1.7B_4B_with_gold_labels_kl_estimation 17 downloads
- RLAIF/iGSM-1M-retry0.1 17 downloads
- RLAIF/dpo_thinking_binary_ultra_feedback_0.02_step_120_with_gold_labels_kl_estimation 16 downloads
- RLAIF/dpo_answer_openorca_offtheshelf_improved_1e-6_0.02_1.7B_0.6B_with_gold_labels_kl_estimation 16 downloads
- RLAIF/dpo_uf_rejudged_mixed_openorca_kl_est 16 downloads
- RLAIF/dpo_answer_only_with_gold_labels_kl_estimation 16 downloads
- RLAIF/train-grm 16 downloads
- RLAIF/dpo_answer_reddit_offtheshelf_1e-6_0.02_4B_4B_with_gold_labels_kl_estimation 15 downloads
- RLAIF/dpo_answer_angel_base_nathan_judged_1e-6_0.02_1.7B_4B_with_gold_labels_kl_estimation 15 downloads
- RLAIF/genrm-ultrafeedback-full-judged-qwen3-4b-base-t07-56989-20250724 15 downloads
- RLAIF/dpo_answer_openorca_offtheshelf_improved_1e-6_0.02_1.7B_4B_with_gold_labels_kl_estimation 14 downloads
- RLAIF/dpo_answer_ultrafeedback_filtered_openorca_1e-6_0.02_0.6B_0.6B_with_gold_labels_kl_estimation 14 downloads
- RLAIF/dpo_thinking_openorca_offtheshelf_improved_1e-6_0.02_1.7B_0.6B_with_gold_labels_kl_estimation 13 downloads
- RLAIF/dpo_answer_reddit_judge_1e-6_0.02_4B_4B_with_gold_labels_kl_estimation 13 downloads
- RLAIF/gm_toy_example 13 downloads
- RLAIF/dpo_thinking_0.05_with_gold_labels_kl_estimation 13 downloads
- RLAIF/dpo_answer_openorca_ppe_improved_1e-6_0.02_1.7B_4B_with_gold_labels_kl_estimation 13 downloads
- RLAIF/WritingPrompts_preferences_chris_filtered 13 downloads
- RLAIF/dpo_thinking_0.02_step_0_with_gold_labels_kl_estimation 13 downloads
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.