Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Published
Source
arXiv
Paper number
1058
Field
Computer Vision
arXiv ID
2609.00111

Key points

  • Without changing the architecture of the pretrained VLM (Qwen3.5-4B), it added just two external modules, a BEV perception head and a Planning Expert, to integrate 3D perception, question answering, and path planning.
  • It used the BEV head as a '3D probe' to make it possible to test how well the model understands actual 3D scene structure, beyond simply producing fluent text answers.
  • Staged training that mixes driving data with general-purpose vision-language data preserved both driving and general capabilities, mitigating catastrophic forgetting (losing old capabilities while learning new ones).
  • It achieved 43.95 mAP for 3D perception and 60.99 mIoU for mapping on nuScenes, a NAVSIM driving score (PDMS) of 90.7, and a Waymo E2E evaluator score of 7.91.
  • In line with the trend toward vehicle cockpits (multimedia) and driving functions sharing a single computing platform, it aims for a model that handles them together rather than a model dedicated only to driving.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)