Robot-Factored World Models via Robot Rendering

Published
Source
arXiv
Paper number
723
Field
Robotics
arXiv ID
2607.22535

Key points

  • It converts actions into nominal trajectories for a robot controller and renders them as input, passing action information without leaking future states.
  • By adding end-effector depth and scene depth, it enables contact and occlusion judgments that are impossible on the 2D image plane.
  • Across both SVD and Wan backbones, it improves PSNR, SSIM, and LPIPS over vector-conditioned baselines.
  • A new robot, xArm6 plus Inspire F1, not used in training can be applied zero-shot simply by changing the rendering.
  • Rendering human manipulation videos after retargeting the hand motions to the robot makes robot manipulation videos possible.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)