Vision as Unified Multimodal Generation
- Published
- Source
- arXiv
- Paper number
- 576
- Field
- Computer Vision
- arXiv ID
- 2607.06560
Key points
- A single UMM unifies structured visual understanding, dense geometry prediction, segmentation, and multi-view 3D, without task-specific heads.
- Text generation handles symbolic outputs, such as detection, OCR, keypoints, and camera pose, while image generation handles dense predictions such as depth, normals, masks, and point maps.
- SenseNova-Vision Corpus is a large training corpus that converts heterogeneous vision annotations into instruction-response examples.
- It outperforms leading systems on structured visual understanding and approaches the best specialized systems on dense geometry and segmentation.
- Task variations can be defined freely in language, including new combinations of color, region, and category that were not seen during training.
- It opens the model and corpus, providing an open-source foundation for reproducible research.
Paper links
External research summaries. These are not HDATF publications or measured product results.