ianlee1996/pokerbench-rl-dpo download history

ianlee1996/pokerbench-rl-dpo is a reinforcement learning dataset on the Hugging Face Hub. In the last 30 days it was downloaded 62 times (13 in the last 7 days), and 177 times in total. It ranks #160,836 among datasets by monthly downloads.

PokerBench RL — Counterfactual DPO Preference Data DPO preference pairs and raw self-play logs for training a Texas Hold'em LLM to exploit non-GTO opponents, addressing the PokerBench paper's Future Work observation that pure SFT models lose to GPT-4-style "donking" strategies. This d

Models trained on pokerbench-rl-dpo

2 models list it as training data.

Open ianlee1996/pokerbench-rl-dpo on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.