Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
- Published
- Source
- arXiv
- Paper number
- 1018
- Field
- Robotics
- arXiv ID
- 2608.26103
Key points
- Because language instructions struggle to convey spatial constraints and intermediate states, it adopted human demonstration videos as the task-instruction interface.
- A pipeline that automatically generates semantically matching human videos from robot trajectories produced HumanGen with 74,200 pairs (8,600 tasks).
- It added an objective for predicting future segments (IFP) to force the model to consult the video prompt rather than rely on shortcuts from existing task knowledge.
- On RoboTwin 2.0, it achieved an average success rate of 47.0% across 7 unseen tasks, a gain of +29.5 percentage points over the strongest baseline (LingBot-VA).
- On real robots, it also followed video instructions well in unseen configurations involving multiple objects, long-horizon tasks, and precision insertion.
Paper links
External research summaries. These are not HDATF publications or measured product results.