SenseNova-U1.5: Towards Native Unified Visual Intelligence
- Published
- Source
- arXiv
- Paper number
- 1080
- Field
- Computer Vision
- arXiv ID
- 2609.11929
Key points
- With a single architecture that drops the encoder and VAE and learns directly from pixels, one model handles image understanding, reasoning, generation, and editing.
- Instead of predicting each pixel patch independently, a reconstruction scheme lets neighboring regions exchange information, producing seamless images even at 4K resolution.
- It uses a specialize-then-unify strategy: separate reinforcement-learning experts are trained per capability and merged into one model via on-policy distillation.
- It scored 68.2 on the VBVR-Pro-Bench text-image bidirectional benchmark, ahead of commercial models such as Nano-Banana-Pro (56.4) and GPT-Image-2 (50.7).
- Despite almost no structured formats in the generation training data, the model follows long, complex structural instructions well, indicating that understanding capability transfers to generation.
- On the reasoning-centric editing benchmark RISEBench, using chain-of-thought substantially improves causal, logical, and temporal editing scores.
Paper links
External research summaries. These are not HDATF publications or measured product results.