Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Published
Source
arXiv
Paper number
949
Field
Robotics
arXiv ID
2608.17512

Key points

  • The Pixel-to-3D Action Formulation lets the VLM operate through familiar 2D visual prompts, avoiding the spatial hallucinations and poor sample efficiency associated with learning 3D geometric transformations.
  • Selective reasoning activates only at critical nodes and compresses the remaining trajectories into space-time indicators, reducing both reasoning latency and memory overflow.
  • Two-level GRPO combines global outcome rewards with fine-grained process rewards to provide dense supervision and teach the model when to trigger reasoning based on the current state.
  • On R2R-CE, it achieved state-of-the-art performance with a 66.2% success rate while requiring only 90,000 training trajectories.
  • When deployed on a physical robot without additional tuning, it achieved a 60.0% success rate, outperforming StreamVLN at 49.0% and DualVLN at 53.0%.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)