Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination
- Published
- Source
- arXiv
- Paper number
- 395
- Field
- Robotics
- arXiv ID
- 2606.10040
Key points
- From WAN-2.2-5B, the authors built a 0.8B compact video expert through structural pruning and knowledge distillation, then combined it with a 0.2B action expert and MoT.
- By lowering the future prediction resolution, they reduced the token count from 240 to 60 and cut latency further from 430 ms to 377 ms.
- Asymmetric denoising with 2 video steps and 10 action steps achieved 139 ms latency with less than 1 percentage point of performance loss.
- It achieved an 86.7 percent average success rate in RoboTwin 2.0 simulation and 66.25 percent success on the real Astribot S1 robot.
- On a consumer RTX 4090, it achieved 98 ms per chunk, which is 32x faster than Motus at 3215 ms and better than pi0.5 at 113 ms.
- The design principle of action-centric future imagination is that the policy needs only structural dynamics cues, not photo-like images.
Paper links
External research summaries. These are not HDATF publications or measured product results.