SenseNova-U1.5: Towards Native Unified Visual Intelligence

Published
Source
arXiv
Paper number
1080
Field
Computer Vision
arXiv ID
2609.11929

Key points

  • With a single architecture that drops the encoder and VAE and learns directly from pixels, one model handles image understanding, reasoning, generation, and editing.
  • Instead of predicting each pixel patch independently, a reconstruction scheme lets neighboring regions exchange information, producing seamless images even at 4K resolution.
  • It uses a specialize-then-unify strategy: separate reinforcement-learning experts are trained per capability and merged into one model via on-policy distillation.
  • It scored 68.2 on the VBVR-Pro-Bench text-image bidirectional benchmark, ahead of commercial models such as Nano-Banana-Pro (56.4) and GPT-Image-2 (50.7).
  • Despite almost no structured formats in the generation training data, the model follows long, complex structural instructions well, indicating that understanding capability transfers to generation.
  • On the reasoning-centric editing benchmark RISEBench, using chain-of-thought substantially improves causal, logical, and temporal editing scores.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)