Vision as Unified Multimodal Generation

Published
Source
arXiv
Paper number
576
Field
Computer Vision
arXiv ID
2607.06560

Key points

  • A single UMM unifies structured visual understanding, dense geometry prediction, segmentation, and multi-view 3D, without task-specific heads.
  • Text generation handles symbolic outputs, such as detection, OCR, keypoints, and camera pose, while image generation handles dense predictions such as depth, normals, masks, and point maps.
  • SenseNova-Vision Corpus is a large training corpus that converts heterogeneous vision annotations into instruction-response examples.
  • It outperforms leading systems on structured visual understanding and approaches the best specialized systems on dense geometry and segmentation.
  • Task variations can be defined freely in language, including new combinations of color, region, and category that were not seen during training.
  • It opens the model and corpus, providing an open-source foundation for reproducible research.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)