PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

Published
Source
arXiv
Paper number
902
Field
Computer Vision
arXiv ID
2608.13552

Key points

  • It enables fair comparison across models by using shared goals and adaptive adjustments from agent players, such as Keep, Stop, Extend, Correct, and End, instead of fixed action trajectories.
  • It evaluates four core dimensions, geometric consistency, interaction fidelity, out-of-view evolution, and insight evolution, along with baseline capability metrics.
  • Among nine models, Genie 3 is the overall best and HappyOyster is second.
  • All models score lowest on out-of-view state evolution and insight evolution, confirming that long-horizon persistence is the current bottleneck.
  • The automatic VQA metric rankings are positively correlated with human preferences.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)