WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

Published
Source
arXiv
Paper number
965
Field
Computer Vision
arXiv ID
2608.20336

Key points

  • It built a unified model that generates group photos of 5–10 reference individuals in one pass. Existing publicly available techniques were limited to a maximum of 4 people.
  • Layout CoT, which plans each person's position, and a location-grounded identity loss (LG-ID) prevent face matching from breaking down as the number of people increases.
  • It surpassed GPT-Image 2 in target-face similarity, 0.499 versus 0.462, and reduced copy-paste artifacts threefold, from 0.169 to 0.055.
  • With 97.3% coverage of requested people and a 2.8% duplication rate, it delivered the most balanced results among the compared systems.
  • It isolated the planning stage, rather than image generation, as the source of remaining errors, indicating the direction for further improvements.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)