Current World Models Lack a Persistent State Core
- Published
- Source
- arXiv
- Paper number
- 462
- Field
- Computer Vision
- arXiv ID
- 2606.20545
Key points
- WRBench defines camera motion as an intervention on observability and evaluates it through a 6-dimensional hierarchical diagnostic chain, from requested-camera to visual integrity to re-observed consistency.
- Natural-25 uses a design with 25 scene families and four event stages to separate spatial movement from state change orthogonally.
- Across 23 models and four control paradigms, the common failure is to restore only the state from the time the object was left when the camera returns, without reflecting events that happened while it was unobserved.
- Scaling Wan from 1.3B to 14B actually worsened re-observed state from 0.66 to 0.62, showing that scaling alone does not solve the problem.
- LingBot-World has the strongest visible-state consistency, at 0.874 and 0.719, but the lowest requested-camera precision at 0.468, revealing a tradeoff between controllability and state consistency.
- The key diagnosis is that all public models currently lack what-memory, meaning a record of hidden changes, and an endpoint-persistence training objective.
Paper links
External research summaries. These are not HDATF publications or measured product results.