Robot-Factored World Models via Robot Rendering
- Published
- Source
- arXiv
- Paper number
- 723
- Field
- Robotics
- arXiv ID
- 2607.22535
Key points
- It converts actions into nominal trajectories for a robot controller and renders them as input, passing action information without leaking future states.
- By adding end-effector depth and scene depth, it enables contact and occlusion judgments that are impossible on the 2D image plane.
- Across both SVD and Wan backbones, it improves PSNR, SSIM, and LPIPS over vector-conditioned baselines.
- A new robot, xArm6 plus Inspire F1, not used in training can be applied zero-shot simply by changing the rendering.
- Rendering human manipulation videos after retargeting the hand motions to the robot makes robot manipulation videos possible.
Paper links
External research summaries. These are not HDATF publications or measured product results.