llm-jp/llm-jp-4-32b-a3b-thinking-dpo-data download history
llm-jp/llm-jp-4-32b-a3b-thinking-dpo-data is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 440 times (80 in the last 7 days), and 6,116 times in total. It ranks #36,608 among datasets by monthly downloads.
llm-jp-4-32b-a3b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4-32b-a3b-thinking. It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (r
Open llm-jp/llm-jp-4-32b-a3b-thinking-dpo-data on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.