ankner/Llama3-8b-ultra-oracle download history

ankner/Llama3-8b-ultra-oracle is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 71 times (25 in the last 7 days), and 1,404 times in total. It ranks #146,423 among datasets by monthly downloads.

Critique-out-Loud Reward Models (CLoud) | Paper | Tweet | Introduction Critique-out-Loud reward models are reward models that can reason explicitly about the quality of an input through producing Chain-of-Thought like critiques of an input before predicting a reward. In c

Open ankner/Llama3-8b-ultra-oracle on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.