jspaulsen/halluci-mate-v1b-dpo download history
jspaulsen/halluci-mate-v1b-dpo is a reinforcement learning dataset on the Hugging Face Hub. In the last 30 days it was downloaded 12 times (1 in the last 7 days), and 121 times in total. It ranks #515,540 among datasets by monthly downloads.
halluci-mate v1b DPO pairs Preference pairs for Direct Preference Optimization fine-tuning of jspaulsen/halluci-mate-v1b, a Qwen3-0.6B chess LLM trained from scratch on UCI moves. Provenance Source: ~11,000 games of jspaulsen/halluci-mate-v1b vs. Stockfish (skill 5, depth 12),
Models trained on halluci-mate-v1b-dpo
1 models list it as training data.
- jspaulsen/halluci-mate-v1c 21 downloads in 30 days
Open jspaulsen/halluci-mate-v1b-dpo on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.