Twins: Learn to Predict Unified Representations with Focal Loss
- Published
- Source
- arXiv
- Paper number
- 722
- Field
- Computer Vision
- arXiv ID
- 2607.22531
Key points
- It builds unified representations by simply concatenating ViT semantic features and VAE detail features channel by channel, with fixed token length.
- It finds that when training a DiT, the ViT branch learns well but the VAE branch does not, and identifies three causes: frequency, dimensionality, and conditional dependence.
- Applying focal loss to flow matching balances the VAE channels by weighting their errors, improving ImageNet gFID by up to 10.57 over MSE.
- Compared with a single encoder such as SigLIP2, it preserves understanding performance while substantially improving reconstruction quality, as measured by PSNR.
- With classifier-free guidance, it achieves an FID of 1.59, which is a solid generation quality result.
Paper links
External research summaries. These are not HDATF publications or measured product results.