Scaling Properties of Text Conditioning in Visual Generation

Published
Source
arXiv
Paper number
786
Field
Computer Vision
arXiv ID
2607.29679

Key points

  • The paper proves that natural-language prompts stop improving image quality as they get longer because they do not add more information, whereas structured prompts continue to add information and improve quality linearly.
  • It introduces two information measures, GPG for image-text alignment and ED for attribute precision, and uses them to quantify the relationship with diffusion loss mathematically.
  • It trains an LLM prompter with SFT, cold start, and verifier-based reinforcement learning so that short user requests are automatically converted into information-rich structured prompts.
  • The final system outperforms every open-source model evaluated on compositional, reasoning, and world-knowledge benchmarks, and it matches or exceeds closed top-tier models.
  • Fine-grained editing, such as object moves, material changes, and scene replacement, becomes possible through field edits in the structured prompt.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)