DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

Published
Source
arXiv
Paper number
739
Field
Robotics
arXiv ID
2607.24159

Key points

  • The method separates video prediction and action prediction into different experts so that policy learning becomes more manageable.
  • It passes representations from multiple layers of the video backbone to the action expert through cross-attention and bridge tokens.
  • It injects auxiliary supervision for object-manipulation regions and relative depth, which strengthens physical information.
  • With 50 demonstrations per RoboCasa task, it reaches a 72.0 percent success rate and improves by up to 22 points over the same budget.
  • On three real-world dual-arm robot tasks, it averages 74 percent and outperforms GR00T-N1.6 at 48 percent and Cosmos Policy at 34 percent.
  • It converges faster with up to 20 times fewer training examples than Cosmos-Policy.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)