Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

Published
Source
arXiv
Paper number
447
Field
Robotics
arXiv ID
2606.18112

Key points

  • It uses a two-axis parameter interface, with task mode and observation parameters such as token budget and camera weight, to support diverse navigation tasks in a single model.
  • Without changing the architecture, randomizing only the training-time parameters makes the Qwen3-VL backbone robust to arbitrary inference-time settings.
  • Training on 15.6M samples and co-training with vision-language data prevents collapse into a reactive action-sequence mapper.
  • It achieves new state-of-the-art results on VLN-CE RxR at 76.5%, EVT-Bench at 90.0%, and NAVSIM at 91.4 PDMS.
  • In agentic system composition, it improves HM-EQA by 10.8% and EXPRESS-Bench by 15.4%, while reducing navigation steps by 77%.
  • It scales favorably from 2B to 8B and shows strong zero-shot generalization in real robot environments.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)