ianlee1996/pokerbench-rl-dpo download history
ianlee1996/pokerbench-rl-dpo is a reinforcement learning dataset on the Hugging Face Hub. In the last 30 days it was downloaded 62 times (13 in the last 7 days), and 177 times in total. It ranks #160,836 among datasets by monthly downloads.
PokerBench RL — Counterfactual DPO Preference Data DPO preference pairs and raw self-play logs for training a Texas Hold'em LLM to exploit non-GTO opponents, addressing the PokerBench paper's Future Work observation that pure SFT models lose to GPT-4-style "donking" strategies. This d
Models trained on pokerbench-rl-dpo
2 models list it as training data.
- ianlee1996/pokerbench-qwen3-14b-lora-dpo-v3 55 downloads in 30 days
- ianlee1996/pokerbench-qwen3-14b-lora-dpo 30 downloads in 30 days
Open ianlee1996/pokerbench-rl-dpo on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.