FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Published
Source
arXiv
Paper number
1127
Field
Computer Vision
arXiv ID
2609.31620

Key points

  • Instead of hand-picking encoder layers, it solves the problem with a single rule (FuseReg): train on random subsets of encoder layers each time.
  • It also showed theoretically that this rule acts as a penalty against sensitivity to disagreements between layers.
  • A single decoder handles full, sparse, and single-layer inputs without retraining, achieving higher reconstruction quality (PSNR) than decoders specialized to fixed fusions.
  • Replacing the decoder alone improved image generation quality (gFID) by 27%, and applying the same rule to the generation stage as well improved it by 29%.
  • The pretrained encoder is left completely untouched, so existing model assets can be reused as-is.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)