Representation Forcing for Bottleneck-Free Unified Multimodal Models
- Published
- Source
- arXiv
- Paper number
- 276
- Field
- Computer Vision
- arXiv ID
- 2605.31604
Key points
- The paper proposes Representation Forcing (RF), a technique that closes this gap by making representation prediction an intrinsic capability of the model.
- If this is simply removed, the model must learn both high-level structure and low-level details directly from raw pixels, which creates a quality gap.
- Specifically, RF forces the decoder to autoregressively predict visual representations as intermediate tokens before pixels, and those tokens are then kept in context to guide pixel diffusion within the same backbone.
Paper links
External research summaries. These are not HDATF publications or measured product results.