WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
- Published
- Source
- arXiv
- Paper number
- 236
- Field
- Computer Vision
- arXiv ID
- 2605.25874
Key points
- The benchmark uses a comprehensive evaluation framework built from five dimensions: video quality, setting compliance, interaction compliance, consistency, and physical compliance.
- It covers 289 cases and 1,058 interaction turns across four interaction types: navigation, action, event editing, and viewpoint switching.
- For navigation, it integrates text, 6-DoF pose, and discrete action inputs to support different input interfaces.
- It combines 22 automatic sub-metrics with evaluation from expert vision models and LMMs, and it validates alignment with human judgments.
- Evaluation of 20 SOTA models showed that no single model dominated all dimensions.
- It provides detailed diagnostic insights into each model's strengths, weaknesses, and open problems.
Paper links
External research summaries. These are not HDATF publications or measured product results.