FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
- Published
- Source
- arXiv
- Paper number
- 1127
- Field
- Computer Vision
- arXiv ID
- 2609.31620
Key points
- Instead of hand-picking encoder layers, it solves the problem with a single rule (FuseReg): train on random subsets of encoder layers each time.
- It also showed theoretically that this rule acts as a penalty against sensitivity to disagreements between layers.
- A single decoder handles full, sparse, and single-layer inputs without retraining, achieving higher reconstruction quality (PSNR) than decoders specialized to fixed fusions.
- Replacing the decoder alone improved image generation quality (gFID) by 27%, and applying the same rule to the generation stage as well improved it by 29%.
- The pretrained encoder is left completely untouched, so existing model assets can be reused as-is.
Paper links
External research summaries. These are not HDATF publications or measured product results.