dRAE: Representation Autoencoder with Hyper-Spherical Codes

Published
Source
arXiv
Paper number
726
Field
Computer Vision
arXiv ID
2607.22148

Key points

  • It clearly diagnoses the cause of VQ codebook collapse as a geometric mismatch between Euclidean distance and the geometry of high-dimensional representation spaces.
  • The best choice is a combination of angle-based, meaning cosine similarity, code assignment and Euclidean commitment loss, while using angles for everything hurts performance.
  • It keeps 100 percent codebook utilization even when the codebook scales to 131,072 entries, and it shows faster convergence and more stable training than VQ and IBQ.
  • On class-conditional ImageNet generation, it reaches gFID 4.45 with HSQ 65536, which is a large improvement over VQ, and it also performs well on text-to-image.
  • With only 12 million image-text pairs, it reaches GenEval 0.63 and DPG-Bench 80.58, which shows an efficient generation pipeline.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)