BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

Published
Source
arXiv
Paper number
823
Field
Robotics
arXiv ID
2608.05042

Key points

  • It is the first framework to integrate temporal memory, which remembers task order, and spatial memory, which remembers the location of occluded objects, into a 3D VLA model.
  • It maximizes data efficiency by converting point clouds into multi-view images for the VLM and predicting actions through heatmaps.
  • On memory-dependent benchmarks such as RMBench and MemoryBench, it achieves an average jump from 20 percent to 93.3 percent compared with the previous state of the art.
  • On real robots such as Franka and Dobot, it shows strong generalization with fewer than 10 training demonstrations and extends to dual-arm manipulation.
  • The heatmap intermediate representation, 2D pretraining, and input alignment are the main contributors to performance, while 3D position input actually hurts performance.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)