GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
- Published
- Source
- arXiv
- Paper number
- 838
- Field
- Robotics
- arXiv ID
- 2608.06332
Key points
- We improved spatial accuracy by using a visual action representation based on URDF rendering instead of numeric actions.
- By separating robot kinematics from environment dynamics, we reduced scene overfitting and improved generalization.
- Autoregressive video prediction supports closed-loop interaction at 8 Hz on an H20 GPU.
- With only limited real data, it achieves robust zero-shot generalization in OOD environments.
- Policies trained on trajectories synthesized by the world model improve overall success from 40.8% to 69.0%.
Paper links
External research summaries. These are not HDATF publications or measured product results.