danilopeixoto/pandora-rlhf download history

danilopeixoto/pandora-rlhf is a text generation dataset on the Hugging Face Hub. In the last 30 days it was downloaded 56 times (10 in the last 7 days), and 1,167 times in total. It ranks #171,591 among datasets by monthly downloads.

Pandora RLHF A Reinforcement Learning from Human Feedback (RLHF) dataset for Direct Preference Optimization (DPO) fine-tuning of the Pandora Large Language Model (LLM). The dataset is based on the anthropic/hh-rlhf dataset. Copyright and license Copyright (c) 2024, Danilo Peixo

Models trained on pandora-rlhf

3 models list it as training data.

Open danilopeixoto/pandora-rlhf on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.