PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

Published
Source
arXiv
Paper number
515
Field
Computer Vision
arXiv ID
2606.28128

Key points

  • We experimentally identify two sources of physical errors: deformation of moving objects and unnatural spatiotemporal correlations between interacting entities.
  • We design complementary pixel-level trajectory alignment loss based on point tracking and semantic-level relation alignment loss based on a frozen video encoder.
  • We propose a physics region focus strategy that concentrates physical supervision on interaction-critical regions, outperforming uniform application by 1.5 points.
  • On R-Bench, we improve Wan2.2-I2V-A14B by 22.3% and Cosmos3-Nano by 9.2%, corresponding to gains of 7.1% and 3.7% over vanilla finetuning, respectively.
  • In the WorldArena action-planner protocol, closed-loop success rises from 16.0% to 24.0%, and average downstream policy success rises from 68.2% to 72.8%.
  • PF-Cosmos achieves the best overall score on R-Bench, PAI-Bench, and EZS-Bench.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)