Let RGB Be the Language of Vision

Published
Source
arXiv
Paper number
637
Field
Computer Vision
arXiv ID
2607.12450

Key points

  • It converts structured visual information such as masks, depth maps, and poses into RGB images, representing multiple tasks as a common RGB-to-RGB editing problem.
  • A single general-purpose image-editing backbone handles understanding tasks and conditional-generation tasks without adding task-specific encoders or decoders.
  • It showed competitive zero-shot performance without separate fine-tuning on tasks including segmentation, depth estimation, and pose-conditioned generation.
  • By connecting visual understanding and generation through the same input-output format, it can serve as a common interface for general-purpose vision-language systems.
  • A limitation is that performance depends heavily on a web-scale-trained image-editing backbone, and gaps remain relative to dedicated models on some specialized tasks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)