World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
- Published
- Source
- arXiv
- Paper number
- 352
- Field
- Robotics
- arXiv ID
- 2606.05979
Key points
- It unifies WAM's future state prediction and VLA's language reasoning in a single autoregressive transformer.
- It uses a dual structure based on meta-queries, with a World Expert for visual prediction and an Action Expert for action generation.
- The World Expert can be disabled at inference time, which enables real-time control with 40 ms inference.
- Test-time scaling further improves robot control accuracy.
- It achieves state-of-the-art results on RoboTwin2.0 Clean at 92.94% and RMBench at 56.5%.
- It also learns new tasks from cross-embodiment videos without action annotations, which points to a scalable direction for robot learning.
Paper links
External research summaries. These are not HDATF publications or measured product results.