WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation

Published
Source
arXiv
Paper number
236
Field
Computer Vision
arXiv ID
2605.25874

Key points

  • The benchmark uses a comprehensive evaluation framework built from five dimensions: video quality, setting compliance, interaction compliance, consistency, and physical compliance.
  • It covers 289 cases and 1,058 interaction turns across four interaction types: navigation, action, event editing, and viewpoint switching.
  • For navigation, it integrates text, 6-DoF pose, and discrete action inputs to support different input interfaces.
  • It combines 22 automatic sub-metrics with evaluation from expert vision models and LMMs, and it validates alignment with human judgments.
  • Evaluation of 20 SOTA models showed that no single model dominated all dimensions.
  • It provides detailed diagnostic insights into each model's strengths, weaknesses, and open problems.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)