dingshizhe/vgtae-cc12m-latents download history
dingshizhe/vgtae-cc12m-latents is a text to image dataset on the Hugging Face Hub. In the last 30 days it was downloaded 409 times (137 in the last 7 days), and 1,001 times in total. It ranks #38,881 among datasets by monthly downloads.
CC12M latents encoded with VGT-AE (448px) Pre-encoded CC12M images in the VGT-AE latent space, so text-to-image training can skip the encoder entirely. These are not DC-AE latents. VGT-AE is a hybrid codec — a fine-tuned Qwen2.5-VL ViT as the encoder, a DC-AE decoder at sampling time.
Open dingshizhe/vgtae-cc12m-latents on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.