WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
- Published
- Source
- arXiv
- Paper number
- 965
- Field
- Computer Vision
- arXiv ID
- 2608.20336
Key points
- It built a unified model that generates group photos of 5–10 reference individuals in one pass. Existing publicly available techniques were limited to a maximum of 4 people.
- Layout CoT, which plans each person's position, and a location-grounded identity loss (LG-ID) prevent face matching from breaking down as the number of people increases.
- It surpassed GPT-Image 2 in target-face similarity, 0.499 versus 0.462, and reduced copy-paste artifacts threefold, from 0.169 to 0.055.
- With 97.3% coverage of requested people and a 2.8% duplication rate, it delivered the most balanced results among the compared systems.
- It isolated the planning stage, rather than image generation, as the source of remaining errors, indicating the direction for further improvements.
Paper links
External research summaries. These are not HDATF publications or measured product results.