LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
- Published
- Source
- arXiv
- Paper number
- 1054
- Field
- Robotics
- arXiv ID
- 2608.30935
Key points
- A single VLM backbone handled instruction following, open-vocabulary object search, and visual tracking in one model without task-specific prediction heads.
- Dual pointing (an affordance point + an object point) raised average success from 54.7% to 63.1%, and removing it reduced performance on every benchmark.
- A residual vector quantization (RVQ) action tokenizer precisely generated a 10-waypoint trajectory using just 3 tokens.
- It achieved the highest monocular success rate across all 10 public navigation simulations.
- It transferred zero-shot to humanoid, quadruped, aerial, and wheeled robots, and expanding environmental diversity was a more reliable driver of performance than increasing model size (the backbone saturated at 4B and above).
Paper links
External research summaries. These are not HDATF publications or measured product results.