Sudhanshu1985/slm-preference-pairs download history

Sudhanshu1985/slm-preference-pairs is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 29 times (6 in the last 7 days), and 97 times in total. It ranks #281,348 among datasets by monthly downloads.

slm preference pairs (DPO / RLAIF) AI-feedback preference pairs used to align the legal/financial SLMs via DPO and RLAIF. On-policy candidates were sampled from each SFT model and ranked by gpt-4.1-mini into (chosen, rejected) pairs. File Model Pairs 125m_pairs.jsonl slm-12

Open Sudhanshu1985/slm-preference-pairs on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.