Let RGB Be the Language of Vision
- Published
- Source
- arXiv
- Paper number
- 637
- Field
- Computer Vision
- arXiv ID
- 2607.12450
Key points
- It converts structured visual information such as masks, depth maps, and poses into RGB images, representing multiple tasks as a common RGB-to-RGB editing problem.
- A single general-purpose image-editing backbone handles understanding tasks and conditional-generation tasks without adding task-specific encoders or decoders.
- It showed competitive zero-shot performance without separate fine-tuning on tasks including segmentation, depth estimation, and pose-conditioned generation.
- By connecting visual understanding and generation through the same input-output format, it can serve as a common interface for general-purpose vision-language systems.
- A limitation is that performance depends heavily on a web-scale-trained image-editing backbone, and gaps remain relative to dedicated models on some specialized tasks.
Paper links
External research summaries. These are not HDATF publications or measured product results.