BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
- Published
- Source
- arXiv
- Paper number
- 823
- Field
- Robotics
- arXiv ID
- 2608.05042
Key points
- It is the first framework to integrate temporal memory, which remembers task order, and spatial memory, which remembers the location of occluded objects, into a 3D VLA model.
- It maximizes data efficiency by converting point clouds into multi-view images for the VLM and predicting actions through heatmaps.
- On memory-dependent benchmarks such as RMBench and MemoryBench, it achieves an average jump from 20 percent to 93.3 percent compared with the previous state of the art.
- On real robots such as Franka and Dobot, it shows strong generalization with fewer than 10 training demonstrations and extends to dual-arm manipulation.
- The heatmap intermediate representation, 2D pretraining, and input alignment are the main contributors to performance, while 3D position input actually hurts performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.