Patch Policy: Efficient Embodied Control via Dense Visual Representations

Published
Source
arXiv
Paper number
666
Field
Robotics
arXiv ID
2607.18236

Key points

  • The core claim is simple: dense visual features, not a giant generative model, are what visual motor policies really need.
  • The method is a minimal extension. It concatenates all patch tokens from a frozen ViT and uses a block causal mask that lets tokens within a frame see each other while blocking information from the future across frames.
  • The policy head can be swapped. It works with both VQ-BeT and Diffusion Policy, and the encoder is not fine-tuned.
  • The results show a 40% relative improvement over global features on four simulation tasks, Push-T, LIBERO Goal, BlockPush, and Cube, and on three real Franka robot tasks.
  • Efficiency is a major strength. The VQ-BeT-based system runs in 10.99 ms at inference even with DINOv2 patches, compared with 61.71 ms for OpenVLA-OFT, and training takes 6.5 GPU hours versus 16 GPU hours.
  • Reducing the number of patches hurts performance. On Push-T, 256 patches score 0.69, while 64 patches score 0.52 and 1 patch scores 0.48, which shows that spatial resolution is the key to precise manipulation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)