PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
- Published
- Source
- arXiv
- Paper number
- 902
- Field
- Computer Vision
- arXiv ID
- 2608.13552
Key points
- It enables fair comparison across models by using shared goals and adaptive adjustments from agent players, such as Keep, Stop, Extend, Correct, and End, instead of fixed action trajectories.
- It evaluates four core dimensions, geometric consistency, interaction fidelity, out-of-view evolution, and insight evolution, along with baseline capability metrics.
- Among nine models, Genie 3 is the overall best and HappyOyster is second.
- All models score lowest on out-of-view state evolution and insight evolution, confirming that long-horizon persistence is the current bottleneck.
- The automatic VQA metric rankings are positively correlated with human preferences.
Paper links
External research summaries. These are not HDATF publications or measured product results.