Twins: Learn to Predict Unified Representations with Focal Loss

Published
Source
arXiv
Paper number
722
Field
Computer Vision
arXiv ID
2607.22531

Key points

  • It builds unified representations by simply concatenating ViT semantic features and VAE detail features channel by channel, with fixed token length.
  • It finds that when training a DiT, the ViT branch learns well but the VAE branch does not, and identifies three causes: frequency, dimensionality, and conditional dependence.
  • Applying focal loss to flow matching balances the VAE channels by weighting their errors, improving ImageNet gFID by up to 10.57 over MSE.
  • Compared with a single encoder such as SigLIP2, it preserves understanding performance while substantially improving reconstruction quality, as measured by PSNR.
  • With classifier-free guidance, it achieves an FID of 1.59, which is a solid generation quality result.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)