psp-dada/Uni-DPO download history

psp-dada/Uni-DPO is an image text to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 85 times (12 in the last 7 days), and 896 times in total. It ranks #127,810 among datasets by monthly downloads.

Dataset Card for ICLR 2026 | Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs 中文 | English Paper Abstract Direct Preference Optimization (DPO) has emerged as a cornerstone of reinforcement learning from human feedback (RLHF) due to its si

Models trained on Uni-DPO

9 models list it as training data.

Open psp-dada/Uni-DPO on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.