Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
- Published
- Source
- arXiv
- Paper number
- 447
- Field
- Robotics
- arXiv ID
- 2606.18112
Key points
- It uses a two-axis parameter interface, with task mode and observation parameters such as token budget and camera weight, to support diverse navigation tasks in a single model.
- Without changing the architecture, randomizing only the training-time parameters makes the Qwen3-VL backbone robust to arbitrary inference-time settings.
- Training on 15.6M samples and co-training with vision-language data prevents collapse into a reactive action-sequence mapper.
- It achieves new state-of-the-art results on VLN-CE RxR at 76.5%, EVT-Bench at 90.0%, and NAVSIM at 91.4 PDMS.
- In agentic system composition, it improves HM-EQA by 10.8% and EXPRESS-Bench by 15.4%, while reducing navigation steps by 77%.
- It scales favorably from 2B to 8B and shows strong zero-shot generalization in real robot environments.
Paper links
External research summaries. These are not HDATF publications or measured product results.