Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

Published
Source
arXiv
Paper number
632
Field
Computer Vision
arXiv ID
2607.13125

Key points

  • It developed four models, Base, Turbo, Edit, and Edit-Turbo, combining a stronger multimodal encoder with agentic prompt rewriting.
  • It jointly refined data quality, training procedures, and inference-time scaling so that improvements in understanding translated into image generation and editing.
  • Using 208.62 million unique images and a theoretical base-model training cost of approximately $400,000, it matched or exceeded several public models on benchmarks.
  • It released weights, code, and training methods under Apache 2.0, allowing reuse in unified image-generation and editing research with smaller budgets.
  • Although it approached leading closed systems, it did not report outperforming them overall, and the comparisons are limited to the scope of published standard evaluations.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)