World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

Published
Source
arXiv
Paper number
352
Field
Robotics
arXiv ID
2606.05979

Key points

  • It unifies WAM's future state prediction and VLA's language reasoning in a single autoregressive transformer.
  • It uses a dual structure based on meta-queries, with a World Expert for visual prediction and an Action Expert for action generation.
  • The World Expert can be disabled at inference time, which enables real-time control with 40 ms inference.
  • Test-time scaling further improves robot control accuracy.
  • It achieves state-of-the-art results on RoboTwin2.0 Clean at 92.94% and RMBench at 56.5%.
  • It also learns new tasks from cross-embodiment videos without action annotations, which points to a scalable direction for robot learning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)