HaoranLiu/DPO-Qwen3-2B-LiteOS download history

HaoranLiu/DPO-Qwen3-2B-LiteOS is a reinforcement learning dataset on the Hugging Face Hub. In the last 30 days it was downloaded 26 times (4 in the last 7 days), and 108 times in total. It ranks #306,850 among datasets by monthly downloads.

DPO-Qwen3-2B-LiteOS Trajectory-level DPO preference pairs for desktop computer-use agents, built on Lite.OSWorld train.perturb. Chosen trajectories come from a GPT-5.5 teacher and from the student's own successes; rejected trajectories are Qwen3-VL-2B-Instruct rollouts on the same tas

Open HaoranLiu/DPO-Qwen3-2B-LiteOS on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.