TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

Published
Source
arXiv
Paper number
938
Field
Computer Vision
arXiv ID
2608.16765

Key points

  • The benchmark represents requests through four operations: object preservation, attribute disentanglement, attribute application, and scene composition. It uses the same formula structure to generate problems, score individual items, and trace the causes of failure.
  • The dataset contains approximately 1,600 evaluation cases, 631 formula templates, and about 4,000 reference images. Complexity ranges from one to eight operation slots.
  • The evaluation compared four commercial models with five open models. Nano Banana 2 achieved the highest average across the four operations at 0.8205, but scored 0.7384 on attribute disentanglement and 0.7989 on attribute application.
  • When 200 Emu3.5 cases were decomposed into simpler subproblems, 38.5% of all diagnostic outcomes were classified as interference arising when multiple references were composed together. Automated analysis and human failure localization agreed in 82.6% of cases.
  • Gemini-2.5-Pro scores the checklist used for the overall evaluation. When humans rechecked 400 outputs produced by two models across 200 cases, item-level agreement for the four scoring models ranged from 85.4% to 88.4%.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)