Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification
- Published
- Source
- arXiv
- Paper number
- 440
- Field
- Computer Vision
- arXiv ID
- 2606.18249
Key points
- A single bitwise visual tokenizer truly unifies the representation space for understanding and generation for the first time.
- Multi-level feature fusion uses shallow layers for low-level detail and deep layers for high-level semantics at the same time.
- Lookup-free BSQ quantization has the effect of a 2^64 codebook, exponentially expanding the vocabulary without an explicit codebook.
- Parallel bitwise prediction achieves a 32x visual compression ratio and predicts only 256 tokens for 1024x1024 images.
- It reaches SOTA on image generation benchmarks such as GenEval and ImgEdit while staying competitive on multimodal understanding benchmarks.
- It trains efficiently with a total of about 33k GPU hours, using a lightweight design with a 400M visual tokenizer and a 2.5B decoder.
Paper links
External research summaries. These are not HDATF publications or measured product results.