Scaling Properties of Text Conditioning in Visual Generation
- Published
- Source
- arXiv
- Paper number
- 786
- Field
- Computer Vision
- arXiv ID
- 2607.29679
Key points
- The paper proves that natural-language prompts stop improving image quality as they get longer because they do not add more information, whereas structured prompts continue to add information and improve quality linearly.
- It introduces two information measures, GPG for image-text alignment and ED for attribute precision, and uses them to quantify the relationship with diffusion loss mathematically.
- It trains an LLM prompter with SFT, cold start, and verifier-based reinforcement learning so that short user requests are automatically converted into information-rich structured prompts.
- The final system outperforms every open-source model evaluated on compositional, reasoning, and world-knowledge benchmarks, and it matches or exceeds closed top-tier models.
- Fine-grained editing, such as object moves, material changes, and scene replacement, becomes possible through field edits in the structured prompt.
Paper links
External research summaries. These are not HDATF publications or measured product results.