penfever/ablation_exploration_in_rl download history

penfever/ablation_exploration_in_rl is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,104 times (176 in the last 7 days), and 2,392 times in total. It ranks #17,945 among datasets by monthly downloads.

Reinforcement Learning Improves Agentic Software Engineering An ablation study of reinforcement-learning (RL) fine-tuning for agentic software-engineering (SWE) models. Starting from an 8B SFT model, we fine-tune with RL across ~20 configurations — varying the objective, loss normaliz

Open penfever/ablation_exploration_in_rl on Hugging Face