AALF/ultrafeedback_wrpo download history
AALF/ultrafeedback_wrpo is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 72 times (14 in the last 7 days), and 1,228 times in total. It ranks #143,699 among datasets by monthly downloads.
[ICLR2025]Weighted-Reward Preference Optimization for Implicit Model Fusion | 📑 WRPO Paper | 🤗 HuggingFace Repo | 🐱 GitHub Repo | Overview In this work, we introduce Weighted-Reward Preference Optimization (WRPO) for the implicit model fusion of heterogeneous ope
Open AALF/ultrafeedback_wrpo on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.