HarnessEval-W: Agentifying the Evaluation of Visual Worlds

Published
Source
arXiv
Paper number
921
Field
Computer Vision
arXiv ID
2608.16859

Key points

  • The evaluation dimensions are eight in total: rendering, physical observation, navigation switching, intent switching, physical switching, suppression of unnecessary scene changes, revisit consistency, and off-screen changes.
  • The evaluator breaks a question into small measurable problems and collects scene evidence using only the skills and tools that are needed.
  • On the same 330 cases, Seedance 2.0 ranks first with an overall score of 75.5, but because the models use different input formats, this should not be treated as a direct apples-to-apples comparison.
  • Compared with 5,000 human pairwise judgments gathered from nine models, the Spearman rank correlation is 0.93 for intent switching and 0.87 for physical switching.
  • With the same video and GPT-5.5 evaluator settings, agreement with human judgment is higher than in the previous WBench, but the validation focuses on only two of eight conditions and may not find the right method for new situations that the evaluation techniques are not ready for.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)