Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
- Published
- Source
- arXiv
- Paper number
- 222
- Field
- Robotics
- arXiv ID
- 2605.22283
Key points
- The geometric estimator network, VGGT, estimates the camera pose and rough scene geometry to establish a global 3D coordinate system.
- The time needed to find the target object was reduced by about 40 to 59 percent.
- The robot showed fewer viewpoint corrections, meaning it less often had to readjust after the head camera located the target, which suggests that spatial memory provided sufficiently reliable estimates for the robot to commit to grasping.
Paper links
External research summaries. These are not HDATF publications or measured product results.