LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

Published
Source
arXiv
Paper number
1054
Field
Robotics
arXiv ID
2608.30935

Key points

  • A single VLM backbone handled instruction following, open-vocabulary object search, and visual tracking in one model without task-specific prediction heads.
  • Dual pointing (an affordance point + an object point) raised average success from 54.7% to 63.1%, and removing it reduced performance on every benchmark.
  • A residual vector quantization (RVQ) action tokenizer precisely generated a 10-waypoint trajectory using just 3 tokens.
  • It achieved the highest monocular success rate across all 10 public navigation simulations.
  • It transferred zero-shot to humanoid, quadruped, aerial, and wheeled robots, and expanding environmental diversity was a more reliable driver of performance than increasing model size (the backbone saturated at 4B and above).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)