DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
- Published
- Source
- arXiv
- Paper number
- 739
- Field
- Robotics
- arXiv ID
- 2607.24159
Key points
- The method separates video prediction and action prediction into different experts so that policy learning becomes more manageable.
- It passes representations from multiple layers of the video backbone to the action expert through cross-attention and bridge tokens.
- It injects auxiliary supervision for object-manipulation regions and relative depth, which strengthens physical information.
- With 50 demonstrations per RoboCasa task, it reaches a 72.0 percent success rate and improves by up to 22 points over the same budget.
- On three real-world dual-arm robot tasks, it averages 74 percent and outperforms GR00T-N1.6 at 48 percent and Cosmos Policy at 34 percent.
- It converges faster with up to 20 times fewer training examples than Cosmos-Policy.
Paper links
External research summaries. These are not HDATF publications or measured product results.