ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
- Published
- Source
- arXiv
- Paper number
- 827
- Field
- Computer Vision
- arXiv ID
- 2608.04436
Key points
- It is the first fully agentic image generation model that unifies reasoning, external tools such as text and image search, and image generation under one UMM policy.
- It develops an optical transformation method for SFT that hides the image-generation tool and leaves only the final image, forcing the UMM to draw directly.
- It uses RAD-GRPO, which combines intent reward and quality reward, to reinforce the full reasoning-search-drawing trajectory.
- With only 7,132 high-quality SFT trajectories, it outperforms partial agentic methods on WISE and WorldGenBench-Humanities.
- It openly releases the training data and the full post-processing infrastructure.
Paper links
External research summaries. These are not HDATF publications or measured product results.