Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
- Published
- Source
- arXiv
- Paper number
- 632
- Field
- Computer Vision
- arXiv ID
- 2607.13125
Key points
- It developed four models, Base, Turbo, Edit, and Edit-Turbo, combining a stronger multimodal encoder with agentic prompt rewriting.
- It jointly refined data quality, training procedures, and inference-time scaling so that improvements in understanding translated into image generation and editing.
- Using 208.62 million unique images and a theoretical base-model training cost of approximately $400,000, it matched or exceeded several public models on benchmarks.
- It released weights, code, and training methods under Apache 2.0, allowing reuse in unified image-generation and editing research with smaller budgets.
- Although it approached leading closed systems, it did not report outperforming them overall, and the comparisons are limited to the scope of published standard evaluations.
Paper links
External research summaries. These are not HDATF publications or measured product results.