HarnessEval-W: Agentifying the Evaluation of Visual Worlds
- Published
- Source
- arXiv
- Paper number
- 921
- Field
- Computer Vision
- arXiv ID
- 2608.16859
Key points
- The evaluation dimensions are eight in total: rendering, physical observation, navigation switching, intent switching, physical switching, suppression of unnecessary scene changes, revisit consistency, and off-screen changes.
- The evaluator breaks a question into small measurable problems and collects scene evidence using only the skills and tools that are needed.
- On the same 330 cases, Seedance 2.0 ranks first with an overall score of 75.5, but because the models use different input formats, this should not be treated as a direct apples-to-apples comparison.
- Compared with 5,000 human pairwise judgments gathered from nine models, the Spearman rank correlation is 0.93 for intent switching and 0.87 for physical switching.
- With the same video and GPT-5.5 evaluator settings, agreement with human judgment is higher than in the previous WBench, but the validation focuses on only two of eight conditions and may not find the right method for new situations that the evaluation techniques are not ready for.
Paper links
External research summaries. These are not HDATF publications or measured product results.