ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Published
Source
arXiv
Paper number
827
Field
Computer Vision
arXiv ID
2608.04436

Key points

  • It is the first fully agentic image generation model that unifies reasoning, external tools such as text and image search, and image generation under one UMM policy.
  • It develops an optical transformation method for SFT that hides the image-generation tool and leaves only the final image, forcing the UMM to draw directly.
  • It uses RAD-GRPO, which combines intent reward and quality reward, to reinforce the full reasoning-search-drawing trajectory.
  • With only 7,132 high-quality SFT trajectories, it outperforms partial agentic methods on WISE and WorldGenBench-Humanities.
  • It openly releases the training data and the full post-processing infrastructure.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)