ankner/Llama3-8b-ultra-self-gen-8b download history

ankner/Llama3-8b-ultra-self-gen-8b is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 48 times (18 in the last 7 days), and 996 times in total. It ranks #191,560 among datasets by monthly downloads.

Critique-out-Loud Reward Models (CLoud) | Paper | Tweet | Introduction Critique-out-Loud reward models are reward models that can reason explicitly about the quality of an input through producing Chain-of-Thought like critiques of an input before predicting a reward. In c

Open ankner/Llama3-8b-ultra-self-gen-8b on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.