Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

Published
Source
arXiv
Paper number
440
Field
Computer Vision
arXiv ID
2606.18249

Key points

  • A single bitwise visual tokenizer truly unifies the representation space for understanding and generation for the first time.
  • Multi-level feature fusion uses shallow layers for low-level detail and deep layers for high-level semantics at the same time.
  • Lookup-free BSQ quantization has the effect of a 2^64 codebook, exponentially expanding the vocabulary without an explicit codebook.
  • Parallel bitwise prediction achieves a 32x visual compression ratio and predicts only 256 tokens for 1024x1024 images.
  • It reaches SOTA on image generation benchmarks such as GenEval and ImgEdit while staying competitive on multimodal understanding benchmarks.
  • It trains efficiently with a total of about 33k GPU hours, using a lightweight design with a 400M visual tokenizer and a 2.5B decoder.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)