GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

Published
Source
arXiv
Paper number
838
Field
Robotics
arXiv ID
2608.06332

Key points

  • We improved spatial accuracy by using a visual action representation based on URDF rendering instead of numeric actions.
  • By separating robot kinematics from environment dynamics, we reduced scene overfitting and improved generalization.
  • Autoregressive video prediction supports closed-loop interaction at 8 Hz on an H20 GPU.
  • With only limited real data, it achieves robust zero-shot generalization in OOD environments.
  • Policies trained on trajectories synthesized by the world model improve overall success from 40.8% to 69.0%.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)