WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
- Published
- Source
- arXiv
- Paper number
- 797
- Field
- Computer Vision
- arXiv ID
- 2608.02603
Key points
- The benchmark spans three manipulation styles, four diagnosis stages, eight evaluation tasks, 1,474 test cases, and 20 evaluated models.
- The eight tasks are Camera Control, Subject Control, Scene Revisit, Terrain Interaction, Object Interaction, Social Interaction, Physical Reaction, and Goal Completion. The first three test instruction following and spatial consistency. The next four test whether the world reacts on its own after seeing the scene, and the final Goal Completion task tests whether the model can plan its own execution steps from only a high-level goal.
- It breaks the same behavior into atomic control units and converts them to the input format each model expects. Camera-based models receive SE(3) camera trajectories, action-based models receive discrete action sequences, and language-based models receive natural-language prompts.
- In World Reactivity cases, the expected reaction is removed from the instruction. For example, the model is told only to move forward toward a staircase, and it must infer the vertical adaptation needed to climb the stairs from the scene itself.
- To match model capabilities, the benchmark is split into two tracks. The static-scene track, where only the camera moves, evaluates all three modalities, while the dynamic-interaction track, where the subject and environment must interact, evaluates only the action and language modalities.
- From the Table 1 comparison, existing benchmarks have many cases but narrow task coverage. WorldScore has 3,000 cases and one task, iWorld-Bench has 4,900 cases and two tasks, while WorldExam covers all eight tasks with only 1,474 cases.
- The conclusion is that high image quality and good instruction following do not imply world reactivity. No model combined broad task coverage with consistently high performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.