Patch Policy: Efficient Embodied Control via Dense Visual Representations
- Published
- Source
- arXiv
- Paper number
- 666
- Field
- Robotics
- arXiv ID
- 2607.18236
Key points
- The core claim is simple: dense visual features, not a giant generative model, are what visual motor policies really need.
- The method is a minimal extension. It concatenates all patch tokens from a frozen ViT and uses a block causal mask that lets tokens within a frame see each other while blocking information from the future across frames.
- The policy head can be swapped. It works with both VQ-BeT and Diffusion Policy, and the encoder is not fine-tuned.
- The results show a 40% relative improvement over global features on four simulation tasks, Push-T, LIBERO Goal, BlockPush, and Cube, and on three real Franka robot tasks.
- Efficiency is a major strength. The VQ-BeT-based system runs in 10.99 ms at inference even with DINOv2 patches, compared with 61.71 ms for OpenVLA-OFT, and training takes 6.5 GPU hours versus 16 GPU hours.
- Reducing the number of patches hurts performance. On Push-T, 256 patches score 0.69, while 64 patches score 0.52 and 1 patch scores 0.48, which shows that spatial resolution is the key to precise manipulation.
Paper links
External research summaries. These are not HDATF publications or measured product results.