Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

Published
Source
arXiv
Paper number
395
Field
Robotics
arXiv ID
2606.10040

Key points

  • From WAN-2.2-5B, the authors built a 0.8B compact video expert through structural pruning and knowledge distillation, then combined it with a 0.2B action expert and MoT.
  • By lowering the future prediction resolution, they reduced the token count from 240 to 60 and cut latency further from 430 ms to 377 ms.
  • Asymmetric denoising with 2 video steps and 10 action steps achieved 139 ms latency with less than 1 percentage point of performance loss.
  • It achieved an 86.7 percent average success rate in RoboTwin 2.0 simulation and 66.25 percent success on the real Astribot S1 robot.
  • On a consumer RTX 4090, it achieved 98 ms per chunk, which is 32x faster than Motus at 3215 ms and better than pi0.5 at 113 ms.
  • The design principle of action-centric future imagination is that the policy needs only structural dynamics cues, not photo-like images.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)