Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

Published
Source
arXiv
Paper number
222
Field
Robotics
arXiv ID
2605.22283

Key points

  • The geometric estimator network, VGGT, estimates the camera pose and rough scene geometry to establish a global 3D coordinate system.
  • The time needed to find the target object was reduced by about 40 to 59 percent.
  • The robot showed fewer viewpoint corrections, meaning it less often had to readjust after the head camera located the target, which suggests that spatial memory provided sufficiently reliable estimates for the robot to commit to grasping.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)