Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
- Published
- Source
- arXiv
- Paper number
- 949
- Field
- Robotics
- arXiv ID
- 2608.17512
Key points
- The Pixel-to-3D Action Formulation lets the VLM operate through familiar 2D visual prompts, avoiding the spatial hallucinations and poor sample efficiency associated with learning 3D geometric transformations.
- Selective reasoning activates only at critical nodes and compresses the remaining trajectories into space-time indicators, reducing both reasoning latency and memory overflow.
- Two-level GRPO combines global outcome rewards with fine-grained process rewards to provide dense supervision and teach the model when to trigger reasoning based on the current state.
- On R2R-CE, it achieved state-of-the-art performance with a 66.2% success rate while requiring only 90,000 training trajectories.
- When deployed on a physical robot without additional tuning, it achieved a 60.0% success rate, outperforming StreamVLN at 49.0% and DualVLN at 53.0%.
Paper links
External research summaries. These are not HDATF publications or measured product results.