Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Published
Source
arXiv
Paper number
1018
Field
Robotics
arXiv ID
2608.26103

Key points

  • Because language instructions struggle to convey spatial constraints and intermediate states, it adopted human demonstration videos as the task-instruction interface.
  • A pipeline that automatically generates semantically matching human videos from robot trajectories produced HumanGen with 74,200 pairs (8,600 tasks).
  • It added an objective for predicting future segments (IFP) to force the model to consult the video prompt rather than rely on shortcuts from existing task knowledge.
  • On RoboTwin 2.0, it achieved an average success rate of 47.0% across 7 unseen tasks, a gain of +29.5 percentage points over the strongest baseline (LingBot-VA).
  • On real robots, it also followed video instructions well in unseen configurations involving multiple objects, long-horizon tasks, and precision insertion.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)