debajyotidasgupta/repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts download history
debajyotidasgupta/repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 81 times (10 in the last 7 days), and 414 times in total. It ranks #133,350 among datasets by monthly downloads.
Reproduction: Contextual Rollout Bandits for RLVR (ICML 2026, #985) Independent reproduction of "Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards" (Lu, Wang, Chai, Yin, Lin, Chen, Luo, Zhuang, Ban, Wang) — OpenReview weMYE1B16x, arXiv 2602.08499. Part of t
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.