llm-jp/llm-jp-4-8b-thinking-dpo-data download history
llm-jp/llm-jp-4-8b-thinking-dpo-data is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 500 times (89 in the last 7 days), and 5,371 times in total. It ranks #33,130 among datasets by monthly downloads.
llm-jp-4-8b-thinking-dpo-data Overview This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4-8b-thinking. It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) r
Open llm-jp/llm-jp-4-8b-thinking-dpo-data on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.