Representation Forcing for Bottleneck-Free Unified Multimodal Models

Published
Source
arXiv
Paper number
276
Field
Computer Vision
arXiv ID
2605.31604

Key points

  • The paper proposes Representation Forcing (RF), a technique that closes this gap by making representation prediction an intrinsic capability of the model.
  • If this is simply removed, the model must learn both high-level structure and low-level details directly from raw pixels, which creates a quality gap.
  • Specifically, RF forces the decoder to autoregressively predict visual representations as intermediate tokens before pixels, and those tokens are then kept in context to guide pixel diffusion within the same backbone.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)