llm-jp/llm-jp-4-33b-thinking-dpo-data download history
llm-jp/llm-jp-4-33b-thinking-dpo-data is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 371 times (42 in the last 7 days), and 873 times in total. It ranks #41,685 among datasets by monthly downloads.
llm-jp-4-33b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4-33b-thinking. It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (r
Open llm-jp/llm-jp-4-33b-thinking-dpo-data on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.