TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
- Published
- Source
- arXiv
- Paper number
- 938
- Field
- Computer Vision
- arXiv ID
- 2608.16765
Key points
- The benchmark represents requests through four operations: object preservation, attribute disentanglement, attribute application, and scene composition. It uses the same formula structure to generate problems, score individual items, and trace the causes of failure.
- The dataset contains approximately 1,600 evaluation cases, 631 formula templates, and about 4,000 reference images. Complexity ranges from one to eight operation slots.
- The evaluation compared four commercial models with five open models. Nano Banana 2 achieved the highest average across the four operations at 0.8205, but scored 0.7384 on attribute disentanglement and 0.7989 on attribute application.
- When 200 Emu3.5 cases were decomposed into simpler subproblems, 38.5% of all diagnostic outcomes were classified as interference arising when multiple references were composed together. Automated analysis and human failure localization agreed in 82.6% of cases.
- Gemini-2.5-Pro scores the checklist used for the overall evaluation. When humans rechecked 400 outputs produced by two models across 200 cases, item-level agreement for the four scoring models ranged from 85.4% to 88.4%.
Paper links
External research summaries. These are not HDATF publications or measured product results.